Destructive testing and inter-laboratory work break the standard crossed study. The nested alternative works, but it buys its answer with an assumption most protocols never state.
A crossed Gage R&R requires that every operator measure every part more than once. When the measurement consumes the unit, that is impossible by construction. The nested design substitutes homogeneity for repetition: instead of measuring one part twice, you measure two units believed to be equivalent.
That substitution is what makes the study possible and it is also its weak point. Everything the study concludes about repeatability depends on whether those units really were equivalent — and that is an assumption, not a measurement.
For when a crossed design is the wrong choice in the first place, see why a Gage R&R gets rejected. This note is about what to do next.
The second is easy to miss because nothing about the data looks unusual. It is the ordinary situation in certification and round-robin testing, and it is why the certification-test characterization in the emissions guard band engagement used a nested design: laboratories did not test a common set of units, so a crossed model would have attributed unit differences to the laboratories.
Units are nested within operator. Each operator receives their own set of groups, and within each group sit several units treated as interchangeable. The operator measures each unit once; the several units within a group play the role that repeat measurements play in a crossed study.
Consumption grows quickly, because nothing is reused:
units consumed = operators × groups per operator × units per group
Three operators, ten groups each, three units per group consumes ninety units. The crossed equivalent would have needed ten. That cost is the reason nested studies get undersized, and undersizing them is the reason they produce unstable variance estimates.
The design trade-off is between groups and units per group at a fixed budget. More groups sharpens the estimate of part-to-part variation. More units per group sharpens the estimate of measurement variation. If the question is whether the measurement system is adequate, spend on units per group. If the question is what the process is doing, spend on groups.
Two things, and both matter for how the result should be read.
The operator-by-part interaction is gone. No part is seen by more than one operator, so there is no information about whether operators disagree differently on different parts. In a crossed study a large interaction is one of the most useful findings available — it points at a procedural ambiguity rather than at hardware. A nested study cannot produce that finding at all.
Repeatability is no longer clean. What the analysis reports as repeatability is the combination of true measurement repeatability and any genuine variation among units inside a group. The nested model cannot separate them, because separating them is exactly what the destroyed part would have allowed.
The direction of the error is predictable and worth stating plainly: whenever the homogeneity assumption is imperfect, a nested study overstates measurement variation. A nested result that fails marginally is therefore weaker evidence of an inadequate gauge than a crossed result that fails by the same margin.
The homogeneity assumption says that units grouped together are interchangeable — that any difference between them is measurement error and not a real difference in the thing measured. It replaces the ability to remeasure. It is load-bearing and it is usually untested.
What makes it plausible is sampling discipline rather than statistics. Units within a group should come from the same lot, the same position, the same time, adjacent material where the physical process allows it. Splitting a single homogeneous sample, where that is possible, is better still because it comes closest to true repetition.
What makes it defensible in review is stating it. A protocol that says units within a group are treated as equivalent; equivalence is supported by common lot and adjacent position; this assumption is not testable within the study will survive scrutiny. A protocol silent on the point invites the question at the worst moment, because an auditor who notices an untested assumption carrying that much weight is entitled to ask what happens if it fails.
The nested ANOVA partitions total variation into an operator term, a group-within-operator term, and a residual. Mapped onto measurement systems language: the operator term is reproducibility, the group-within-operator term carries part-to-part variation, and the residual is the compromised repeatability described above.
Two cautions follow. First, apply the acceptance ratio deliberately — percent tolerance and percent study variation answer different questions here exactly as they do in a crossed study, and the nested design does nothing to change that. Second, resist quoting the number of distinct categories from a nested study without comment, because it inherits both the part-selection dependence and the inflated repeatability.
When a nested study fails, the productive question is usually not is the gauge adequate but how much of this is the homogeneity assumption. That is answerable, but by improving the sampling and repeating, not by reanalyzing what you have.
With a nested design, and by substituting homogeneity for repetition. Because the unit is consumed, no part can be measured twice, so repeat measurement is replaced by measuring several units believed to be effectively identical — adjacent pieces from the same lot, position or batch. Each operator or laboratory receives its own set of these units, nested within that operator. The analysis is a nested ANOVA that estimates repeatability and reproducibility but cannot estimate the operator-by-part interaction, because no part is shared. The validity of the whole study rests on whether the units within a group really are equivalent.
The operator-by-part interaction, and a clean repeatability term. The interaction is unestimable because no part is measured by more than one operator, so the two effects are confounded by construction. Repeatability is also compromised: what the analysis labels repeatability is really the combination of true measurement repeatability and any genuine unit-to-unit variation within the supposedly homogeneous group. Those two cannot be separated from the study alone. A nested design therefore tends to overstate measurement variation whenever the homogeneity assumption is imperfect, which it usually is.
The assumption that units grouped together for repeat measurement are interchangeable — that any difference between them is measurement error rather than a real difference in the thing being measured. It replaces the ability to remeasure a single part. It is doing most of the work in the study and it is rarely tested. Sampling adjacent material, from the same lot, the same position and the same time, is how the assumption is made plausible. Stating it explicitly in the protocol is how the study survives review, because an auditor who spots an untested assumption doing that much work will ask about it.
Whenever the same part cannot be measured more than once by more than one operator. Destructive or one-shot testing is the clear case: burst and leak testing, tensile and peel tests, sterility and bioburden work, many assays. The second case is inter-laboratory comparison where laboratories do not test a common set of units, so laboratory and unit effects cannot be separated. Running a crossed analysis in either situation does not fail loudly; it returns ordinary-looking numbers in which part-to-part variation has been absorbed into repeatability.
More than a crossed study, because parts cannot be reused. Total consumption is operators multiplied by groups per operator multiplied by units per group, so three operators, ten groups and three units per group consumes ninety units rather than the ten a crossed study would need. The design trade-off is between groups and units per group: more groups improves the estimate of part variation, more units per group improves the estimate of measurement variation. When the question is whether the measurement system is adequate, favor units per group. When the question is what the process is doing, favor groups.