Device agreement and test–retest reliability

A new instrument is almost never compared against truth. It is compared against another instrument that also carries error — and most of what goes wrong in device validation follows from forgetting that.

Short answer

Three different questions get asked of a validation study, and they need three different analyses:

A study that reports a correlation coefficient has answered none of them.

Why correlation is the wrong statistic

Correlation measures association, not agreement. A device reading exactly twice the reference value correlates perfectly and agrees with nothing at all. Bias of that kind is invisible to the coefficient.

Worse, correlation is inflated by the spread of the sample. Compare two instruments across a wide range of subjects and the coefficient rises regardless of how well they agree on any individual. That is why method-comparison papers reporting r above 0.95 can describe a device nobody should use interchangeably — the number is partly a property of who was recruited.

The validation question is never do these two move together. It is how far apart are they likely to be on the next patient, and that has to be answered in the units a clinician thinks in.

Agreement: bias and limits

The standard approach plots the difference between the two methods against their average, which separates two things a scatterplot against the reference confounds: systematic bias, and how that bias behaves across the measurement range.

It yields the mean difference — the systematic offset — and limits of agreement spanning roughly 95% of individual differences. Bias is often correctable by calibration. Limits usually are not.

The critical step happens before any data are examined: state how large a difference would change clinical management, and write it into the protocol. Only then ask whether the limits fall inside it. Reversing that order — computing limits and then deciding whether they seem acceptable — is how a study reaches whatever conclusion the sponsor needed, and a reviewer will recognise it.

Two patterns worth watching for. Bias that changes across the range means a single correction will not work. And limits computed on too few subjects are themselves imprecise, so the interval around the limits should be reported alongside them.

Reliability is a different question

Agreement is measured in the units of the instrument. Reliability is a ratio — subject variance divided by total variance — and it asks how well the measurement separates people from one another.

That makes the intraclass correlation population-dependent in exactly the way percent study variation is part-dependent in an industrial gauge study. The same instrument scores higher on a heterogeneous sample than a narrow one without changing in any respect. It is the identical statistical trap in clinical vocabulary, and it catches people who would never fall for it on a factory floor.

Two practical requirements. Report which ICC form was used, since the definitions differ in whether raters are treated as fixed or random and whether the intended use is a single measurement or an average — they can produce materially different values on the same data. And pair the ICC with an absolute measure such as the standard error of measurement, from which the smallest detectable change follows. That last quantity is usually the one a clinician actually needs: it says how much a patient's score must move before the movement means anything.

When the reference is also imperfect

Most device validations compare against a predicate instrument or an established instrument, not against truth. Both carry error, so disagreement is jointly owned and cannot be attributed to the new device from a two-method comparison alone.

Three consequences, all worth building into the protocol rather than discovering in review:

This is the same structure as an inter-laboratory measurement study: without repeated observations, instrument effects and subject effects are confounded and no amount of analysis separates them afterwards.

Normative ranges

Establishing whether an individual's value is unusual requires a reference distribution, and a reference distribution built without the covariates that genuinely shift the measurement will misclassify systematically rather than randomly.

For cognitive and functional measures, age and education routinely move the distribution. A single range ignoring them will over-identify impairment in older or less-educated subjects and under-identify it in younger or more-educated ones — a bias that falls on identifiable groups, which makes it a fairness problem as well as a statistical one.

Covariate-adjusted ranges built by regression address this, with two requirements that are frequently missed. The reference sample must be adequate at the extremes, because that is where classification decisions are made and where a regression is least certain. And the interval must be a prediction interval for a new individual, not a confidence interval for the mean — these are very different widths, and using the narrow one produces a range that flags far more people than intended.

The same question in a different vocabulary

Every problem on this page has an industrial twin. Limits of agreement and a guard band both convert measurement uncertainty into a decision rule. The ICC and percent study variation are both ratios that inherit the spread of whoever was sampled. Test–retest reliability is repeatability. An imperfect reference is an inter-laboratory study without shared units.

That is not a coincidence and it is not a metaphor — it is the same variance-components problem, and the statistical machinery transfers directly. The six validation and reliability studies described on the Work page turned on exactly that question: whether an observed difference reflected the patient, the device, or the normative reference.

Send the problem, not the data →

Common questions

Why is correlation the wrong statistic for method comparison?

Because correlation measures association, not agreement. A device reading consistently twice the reference value correlates perfectly and agrees with nothing. Correlation is also inflated by the range of the sample — comparing instruments across a wide spread of subjects produces a high coefficient regardless of how well they agree on any individual, which is why method-comparison papers reporting r above 0.95 can still describe a device nobody should use interchangeably. The question in validation is how far apart two measurements on the same subject are likely to be, and correlation does not answer it.

What do Bland–Altman limits of agreement tell you?

How far apart the two methods are likely to be on an individual subject. The analysis plots the difference between methods against their average, then reports the mean difference — the systematic bias — and limits of agreement covering roughly 95 percent of differences. The clinical judgment comes first and separately: decide before looking at the data how large a difference would change management, then ask whether the limits fall inside it. A device can show negligible bias and still be unusable if the limits are wide, because bias describes the average subject and no one is treating the average subject.

What is the difference between agreement and reliability?

Agreement asks how close two measurements are in the units of the measurement. Reliability asks how well the measurement separates subjects from one another, which is a ratio of subject variance to total variance. The intraclass correlation is a reliability measure, so it inherits a dependence on the population sampled: the same instrument tested on a heterogeneous group posts a higher ICC than on a narrow one, without changing at all. Report both, and report the ICC form used, because the several definitions differ in whether raters are fixed or random and whether single or averaged measurements are intended.

How do you validate a device when the reference is also imperfect?

By stopping the pretence that the comparison establishes accuracy. When the reference carries its own error, disagreement is jointly owned and cannot be attributed to the new device from a two-method comparison alone. Three consequences follow. Do not call the comparator a gold standard in the protocol unless it genuinely is. Frame the claim as interchangeability rather than accuracy. And where possible include repeat measurements on both instruments, because that separates each instrument's own repeatability from the disagreement between them — which is the only way to say whose error is driving the result.

What makes a normative range defensible?

That it accounts for the characteristics which genuinely shift the measurement, and that its uncertainty is reported. Age and education commonly move cognitive measures; ignoring them produces a range that misclassifies older or less-educated subjects systematically. Covariate-adjusted ranges built by regression handle this, but the reference sample has to be large enough at the extremes where classification actually happens, and the interval should be a prediction interval for a new individual rather than a confidence interval for the mean. Those are very different widths and confusing them is a frequent and consequential error.