Why Practice Tests Alone Will Not Get You the Score

There is a preparation pattern so common it deserves a name. A candidate buys a book of official practice tests, works through one every week, marks it, records the score, and repeats. The score climbs for the first two or three tests, then stops. Test five looks like test four. Test eight looks like test five. The candidate concludes they need more practice tests.

The problem is not effort, and it is not the tests. It is what a practice test is actually built to do.

A practice test is a thermometer, not a diagnosis

A practice test measures your level under exam conditions. That is genuinely valuable, and nothing in this post argues against doing them. But measurement and diagnosis are different activities, and the confusion between them is where preparation time disappears.

A thermometer tells you that you have a fever. It does not tell you whether the cause is an infection, an inflammation, or something else entirely, and taking your temperature every hour does not treat any of them. A practice test works the same way: it tells you where you are, with reasonable accuracy, and it says almost nothing about why you are there or what specifically would move you.

Cambridge says this itself, in the same document that publishes the official conversion tables for turning practice test scores into estimated Cambridge Scale scores. The tables, in Cambridge's own words, should not be used to try to predict precise scores in the live exam, but can be a useful diagnostic tool, indicating areas of relative strength and weakness. The instrument was designed for diagnosis. Most candidates use it as a forecast, record the headline number, and throw away the only part that was ever going to help.

What the headline score hides

Suppose a candidate scores the equivalent of 158 on a full B2 First practice test. That number sits just below the 160 boundary, and the natural conclusion is close, keep going. But 158 is an average of five separate scores, and the profiles hiding inside it are not remotely equivalent.

One candidate at 158 is even across all five skills: everything sits between 155 and 161. Their gap is thin and distributed, and general exposure at the right level genuinely might close it.

Another candidate at 158 is at 172 in Reading, 168 in Listening, 165 in Speaking, and 139 in Use of English. Their overall number is nearly identical, and their situation is completely different: four skills are already comfortably above the boundary, and one is far enough below it to cap the entire result. More full practice tests are close to useless for this candidate, because four fifths of every test they sit is rehearsing skills that were never the problem.

The headline score cannot distinguish these two candidates. The five scores underneath it can, which is the entire argument of how the Cambridge Scale actually works: the exam reports one number, but preparation needs five.

Repetition rehearses your errors too

There is a second, quieter problem with the test-mark-repeat loop. Doing a task repeatedly does not improve the task unless something changes between attempts. A candidate who loses marks on key word transformations because of one specific grammatical pattern will lose those marks on every practice test, in the same place, for the same reason, until someone identifies the pattern. The test registers the loss. It never explains it.

This is how candidates arrive at week ten with a folder of completed tests, a flat score line, and the sincere belief that they have been preparing seriously. They have been measuring seriously. The preparing part, the targeted work on the specific mechanism dropping the marks, never actually started, because nothing in the loop was capable of finding it.

The skill posts in this series exist because the mechanisms live at a level of detail the headline score cannot see. The Writing paper scores four separate criteria, and a candidate can lose marks on three of them while drilling the fourth. The Speaking test scores five separate marks across four parts, each part built to surface a different one. A practice test records the output of all of that machinery as a single number per paper. The machinery is where the marks are.

What practice tests are actually for

Used correctly, practice tests have three legitimate jobs, and they do all three well.

The first is calibration at the start: establishing where you genuinely are, per skill, before deciding anything. The second is rehearsal near the end: timing, stamina, and familiarity in the final weeks before the exam date, when the gaps have already been closed and the task is execution. The third is checkpointing in between: a periodic measurement, spaced weeks apart, confirming that the targeted work is moving the specific score it was supposed to move.

What sits between those checkpoints is the actual preparation: work aimed at the identified weakness, not another lap around all five skills. The test is the instrument panel. It was never the engine.

The honest sequence

The sequence that works is unglamorous and specific. Measure once, properly, per skill. Identify which score is holding the average down and why, at the level of task types and error patterns, not skill names. Work on that, and nearly only that. Measure again after the work, not instead of it. This is the entire logic behind a diagnostic-first approach, and it is also why a candidate who knows their five scores can prepare in a fraction of the time of a candidate who knows only their average: they are not spending four fifths of every session on skills that were already above the line.

Frequently asked questions

How many practice tests should I do before B2 First?

Fewer than most candidates do, spaced further apart. One at the start for calibration, one every few weeks as a checkpoint on targeted work, and two or three in the final stretch for timing and stamina. A practice test every week with nothing changing between them measures the same gap repeatedly without closing it.

My practice test scores stopped improving. What does that mean?

It usually means the remaining gap is concentrated in a specific skill or task type that repetition cannot reach. The plateau is not a signal to do more tests. It is a signal to find out, at the level of individual papers and error patterns, exactly where the flat line is coming from.

Are official Cambridge practice tests accurate predictors of my real score?

Cambridge's own guidance says no. Converted practice scores indicate relative strengths and weaknesses, and scores within roughly three scale points of a boundary should be treated as unresolved. They are designed as a diagnostic instrument, not a forecast.

Is it bad to prepare using only practice tests?

It is incomplete rather than bad. Tests measure and rehearse, and both matter. What they cannot do is identify why a specific score is stuck or teach the mechanism that unsticks it, and preparation built only on testing leaves that entire layer missing.

Where this actually starts

Everything above assumes you know your five scores and what sits underneath them, and that is precisely the information most candidates never have. The Fluensys Proficiency Profile exists to produce it: a per-skill measurement referenced against the Cambridge Scale, reviewed by a human instructor, identifying not just which score is holding you back but which specific patterns inside it are doing the holding. A practice test tells you the score is stuck. The diagnostic tells you where, and why, and that difference is the plan.