What an accuracy percentage means

Products get sold on a single number — ninety-eight percent accurate, ninety-nine percent agreement. The number is usually real. What it does not say is what it was measured against, whether it describes bias or spread, and how the worst cases behaved. Those omissions are where the argument starts.

Short answer

An accuracy percentage is a summary of a particular study, not a property of the method. Reported alone it compresses at least four separable things into one figure, and each of them can move the number independently.

trueness · precision · the reference's own error · the level of aggregation

None of these is recoverable from the headline figure. That is not a criticism of anyone's product; it is a statement about what one number can carry. The useful question is never whether the number is honest. It is what the number would have to say in order to be checkable.

Accuracy is two things wearing one word

In ordinary use, accuracy means closeness to the truth. In measurement it splits into trueness — how close results sit to the true value on average — and precision — how tightly repeated measurements cluster around their own mean.

The two are independent, and the failure modes are different in kind:

A headline percentage that merges them tells you a study produced a satisfying average and nothing about which failure mode the method has. Since the remedies are entirely different — recalibration for one, a better measurement process for the other — the distinction is the first thing worth recovering.

The reference is not the truth

Almost every accuracy figure comes from comparing a method against something treated as correct: a benchmark study, a manual measurement, an established instrument. That comparator has error of its own, and the arithmetic of what follows is not subtle.

Let the method and the reference each measure the same true value with their own independent error. The difference between them carries both errors, so the variance of the observed differences is the sum of the two error variances. The disagreement you measure therefore overstates the method's own error, and it overstates it by exactly the reference's error variance.

Which means a method compared against a noisy reference looks worse than it is — and the shortfall is recoverable, but only if somebody has characterized the reference rather than assuming it perfect.

The reverse case is more troubling. If the method was calibrated against the same reference it is later validated against, it inherits that reference's bias and the comparison cannot see it. Agreement is then excellent and trueness is unknown, because both are wrong together. Nothing in the reported percentage distinguishes this from genuine accuracy. Only knowing how the calibration and the validation were separated does.

Regression dilution, and the bias that is not there

A common next step is to regress the new method on the reference and inspect the slope for proportional bias. When the reference carries measurement error, that slope is attenuated toward zero, and the more error it carries relative to the spread of true values, the stronger the attenuation.

What this looks like in a report is a method that appears to under-respond at high values and over-respond at low ones. It is a recognizable, plausible-looking pattern, and it is frequently an artifact of having put an error-laden variable on the x-axis. Correcting the method for it makes the method worse.

Ordinary least squares assumes the predictor is measured without error, which in a method comparison is exactly the assumption that fails. Errors-in-variables approaches such as Deming regression are built for this, and they need an estimate of the ratio of the two error variances — which is another reason the reference has to be characterized rather than trusted.

Percent of what, computed where

Two structural questions decide how a percentage behaves, and neither is visible in the number.

Percent of what. A mean absolute percentage error, the proportion of cases falling within some tolerance, and the correlation between two methods are three different quantities that all get written as a percentage. The proportion-within-tolerance form is the slipperiest, because it depends entirely on a tolerance somebody chose, and widening that tolerance improves the number without changing the method at all.

Computed where. Errors that are independent across components partially cancel when the components are summed. A figure computed on an assembled whole can therefore look substantially better than the per-component figure underneath it, and the whole-object number is usually the one that gets published. Both can be correct. They answer different questions, and only one of them is relevant if a decision turns on a single component.

This is the same aggregation problem that makes a gauge study look acceptable on averages while individual readings are not fit for the decision being made.

The tails are where the argument happens

An average error describes the typical case. Nobody disputes the typical case.

A mean absolute error of two percent is entirely compatible with a small share of cases being wrong by twenty, and those are the cases that become a claim, an appeal or a lawsuit. The mean is the least informative summary for exactly the situations that carry consequence.

Limits of agreement answer the question the average dodges: within what interval do most individual differences fall? That is what someone deciding on one measurement needs, and it is why a method-comparison analysis should present the distribution of differences rather than collapsing it. The approach also has the virtue of not requiring either method to be a gold standard, which is honest when neither is. The same machinery underlies inter-method agreement and test–retest reliability in device validation.

What conditions did the study run under

A validation study is a sample from a population of use conditions, and the number it yields is only as general as that sample.

The recurring gaps are conditions the study did not span: a narrower range of inputs than routine practice sees, favorable environmental conditions, one operator or one site standing in for many, equipment and software at a single version, and cases selected for being measurable rather than sampled to represent the real mix. Each of these produces a number that is true of the study and optimistic about deployment.

Where the sources of variation are structurally different — sites that cannot measure the same item, or material consumed by the test — the study has to be nested rather than crossed, or a whole component of variation gets misattributed and the accuracy figure inherits the mistake.

What a defensible claim looks like

None of the above argues for publishing less. It argues for publishing a claim that survives someone competent reading it adversarially, which is a commercial advantage rather than a concession.

A claim that states these can be argued with. That is the point: a number nobody can interrogate is also a number nobody can defend, and the first serious challenge is a poor moment to discover which one you have.

The related question — what a measurement's error does to a pass or fail decision made from it — is the guard band problem, and what it does to a capability index is covered separately. All of it is the same variance-components question asked from different directions.

Send the problem, not the data →

Common questions

What does a single accuracy percentage actually tell you?

On its own, very little that can be checked. Accuracy is not one quantity — it combines trueness, meaning how close results sit to the true value on average, and precision, meaning how tightly repeated measurements cluster. A method can be badly biased and highly precise, or unbiased and hopelessly scattered, and both can be reported as the same headline percentage. A single number also cannot reveal what it was computed against, at what level results were aggregated, or how the worst cases behaved. It is a summary of a study, not a property of the instrument, and without the study it is not a claim anyone can verify.

Why does comparing a method against an imperfect reference bias the result?

Because the differences you observe contain both methods' errors, not just the one being evaluated. If the new method and the reference each measure the same true value with independent error, the variance of their difference is the sum of the two error variances. The disagreement therefore overstates the new method's own error, and the overstatement equals the reference's error variance. That direction reverses if the method was calibrated against the same reference: it then inherits the reference's bias and looks closer to truth than it is. Both effects are quantifiable, but only if the reference's own uncertainty has been characterized rather than assumed to be zero.

What is regression dilution and why does it matter here?

When a new method is regressed on a reference that carries measurement error, the estimated slope is pulled toward zero. The larger the reference's error relative to the spread of true values, the stronger the pull. In a method comparison this shows up as an apparent proportional bias that does not exist — the new method looks as though it under-responds at high values and over-responds at low ones. Analysts then sometimes correct for a bias that is an artifact of the analysis. Errors-in-variables methods such as Deming regression exist precisely to handle this, and they require an estimate of the error ratio between the two methods.

Why do the tails matter more than the average error?

Because decisions and disputes happen at the extremes, not at the mean. A mean absolute error of two percent is compatible with a small fraction of cases being off by twenty, and it is those cases that generate the challenge, the claim or the appeal. Limits of agreement address this directly: rather than a single average, they state the interval within which most individual differences fall, which is what someone deciding on one measurement actually needs. An average error also tends to shrink as results are aggregated, so a figure computed per assembled object can look far better than the per-component figure underneath it.

What does a defensible accuracy claim have to state?

Six things. What the comparison was made against, and what that reference's own uncertainty is. Whether the figure describes bias or spread, reported separately rather than merged. The unit of analysis, and whether the number was computed before or after aggregation. The distribution of individual differences, usually as limits of agreement, not only the average. The conditions the study ran under, and how far those match routine use. And the sampling scheme, since a convenience sample of easy cases produces a number that does not generalize. A claim stating all six can be checked and argued with, which is exactly what makes it hold up.