A failed study is a diagnosis, not a verdict. Three quite different things produce the same red number, they call for opposite responses, and telling them apart rarely requires collecting anything new.
A rejected Gage R&R usually means one of three things:
The responses are not interchangeable. The first calls for a repeat study with better sampling, the second for reporting the correct ratio, the third for capital. Deciding by reflex — usually "buy a better gauge" — is how a measurement problem becomes an expensive one.
A Gage R&R partitions total observed variation into part-to-part variation and measurement variation, and splits the latter into repeatability (the same operator, same part, repeated) and reproducibility (differences between operators, and any operator-by-part interaction). Those variance components are the output. What turns them into a pass or fail is the denominator you divide by, and there are two defensible choices.
Percent study variation compares measurement variation to the total variation in the parts you sampled. It answers: can this system see the differences between the things I am measuring? That is the right question when the measurement feeds process control — control charts, capability studies, detecting a shift.
Percent tolerance compares measurement variation to the specification width. It answers: can this system decide whether a unit conforms? That is the right question when the measurement supports a release or acceptance decision.
These diverge sharply when process variation is small relative to the tolerance — which is the situation a capable process is supposed to be in. A gauge measuring a tightly controlled process can post an alarming percent study variation while being entirely adequate against the specification. The familiar ten and thirty percent thresholds are meaningless without saying which ratio they apply to, and a study that reports one number without naming the ratio cannot be acted on.
Percent study variation carries part variation in its denominator, so anything that shrinks the spread of the sampled parts inflates the measurement system's apparent share. Parts pulled from a single shift, a single lot, or — most damaging — deliberately selected as representative good units will do exactly that. The gauge has not changed. The denominator has.
The diagnostic is simple. If a study fails on percent study variation but passes comfortably on percent tolerance, look hard at how the parts were chosen before concluding anything about the instrument. Parts for a Gage R&R should span the range the process genuinely produces, tails included, which is not the same as selecting parts that span the specification.
The corollary matters for planning: a Gage R&R done on a stable, capable process will often look worse than the same gauge assessed on a process running wider. That is a property of the ratio, not evidence of degradation.
The standard crossed study has every operator measure every part more than once. That structure is what allows repeatability, reproducibility and the interaction to be separated. When the structure is not achievable, the crossed analysis does not fail loudly — it returns numbers that look ordinary and are wrong.
Two situations require a nested design instead, with units nested within operator or laboratory:
Run crossed anyway and the repeatability estimate quietly absorbs part-to-part variation, which inflates measurement variation and pushes the study toward rejection. A study rejected for that reason is not telling you about your gauge at all. The certification-test characterization described in the emissions guard band engagement used a nested design for precisely this reason: laboratories did not test a common set of units.
Sometimes the gauge genuinely cannot resolve the tolerance. The signature is consistency: the study fails on percent tolerance, fails on percent study variation, and returns a low number of distinct categories, with parts sampled across a defensible range and the correct design used. At that point the useful question shifts from whether the system is adequate to what can be done short of replacing it.
Averaging repeated measurements reduces repeatability by the square root of the number of repeats and is often enough when the shortfall is modest, at a known cost in cycle time. Reproducibility dominating instead points at method and training rather than hardware, and is usually cheaper to fix. A large operator-by-part interaction says operators disagree differently on different parts, which is a procedural ambiguity — a fixturing or judgment step that the written method does not pin down. None of those require capital.
Measurement variation does not stay in the measurement system. It propagates into every capability index computed from the same data, so a process judged incapable may be a capable process observed through a noisy gauge. Separating the two is a variance-components question, and it is answerable from records most manufacturers already hold.
In the first-pass yield investigation described on the Work page, that separation was the point of the exercise: establishing how much of the observed variation belonged to the measurement system before any experiment was designed, so that effort went where the variability actually was rather than where it appeared to be.
Establish which of three things happened before changing anything. The study design may not have matched the measurement process — most often the parts selected did not span the range the process actually produces. The acceptance criterion may have been the wrong one for the application, because percent study variation and percent tolerance answer different questions and a gauge can fail one while passing the other. Or the measurement system genuinely cannot resolve the tolerance. These call for completely different responses, and determining which applies can usually be done from the data already collected rather than by repeating the study.
Percent study variation compares measurement variation to the total variation observed in the parts you sampled. Percent tolerance compares it to the specification width. They answer different questions. If the measurement supports process control — detecting shifts, running control charts, judging capability — percent study variation is the relevant ratio. If it supports a conformance decision against a specification, percent tolerance is. Reporting only one, or applying the ten and thirty percent thresholds without stating which ratio they refer to, is the single most common reason a perfectly adequate gauge gets condemned.
Because percent study variation has part variation in its denominator. If the parts chosen sit close together — pulled from one shift, one lot, or deliberately picked as good units — the denominator shrinks and the measurement system's share inflates, even though the gauge has not changed. The parts must span the range the process genuinely produces, including the tails. A study that fails on percent study variation but passes on percent tolerance is very often a part-selection artifact rather than a gauge problem.
Whenever the same part cannot be measured more than once by more than one operator. Destructive or one-shot testing is the obvious case: the unit is consumed, so repeat measurements are impossible by construction. A subtler case is inter-laboratory work, where laboratories do not test a common set of units, so laboratory and unit effects cannot be separated by a crossed design. Both require a nested design, in which units are nested within operator or laboratory. Running the crossed analysis anyway produces a repeatability estimate that silently contains part-to-part variation.
It estimates how many separate groups the measurement system can reliably distinguish within the observed part variation, computed as roughly 1.41 times the ratio of part variation to measurement variation. Five or more is the conventional minimum. Below two the system is effectively an attribute gauge — it can sort into pass and fail but cannot rank. Because it derives from the same variance components as percent study variation, it carries the same dependence on part selection, and it should never be quoted on its own.