When the measurement is a category rather than a number — pass or fail, defect class, severity grade — variance has nothing to partition and Gage R&R does not apply. The replacement study is easy to run and remarkably easy to run in a way that proves nothing.
A category has no spread, so there is no repeatability variance to estimate. What can be estimated is agreement, and it has to be assessed three separate ways:
A study reporting a single overall agreement figure cannot distinguish these, and the remedies differ completely: retraining for the first, a clearer written standard for the second, a corrected standard or recalibrated judgment for the third.
Two inspectors who both call almost everything good will agree almost all the time without looking at anything. On product that is 98% conforming, two appraisers who simply always say pass agree 98% of the time and have demonstrated nothing whatever about the inspection system.
Percent agreement does not subtract the agreement you would expect by chance. Kappa statistics do, which is why they are the standard reporting measure — Cohen's kappa for two appraisers, Fleiss' kappa for more than two.
κ = (Po − Pe) ÷ (1 − Pe)
where Po is observed agreement and Pe is the agreement expected from the marginal rates alone.
Correcting for chance introduces a problem of its own on exactly the same skewed data. When one category dominates, Pe is already close to one, so the denominator becomes small and even excellent observed agreement collapses to a low kappa.
Two appraisers agreeing 95% of the time on product that is 97% conforming can post a kappa near zero. Reported without context that reads as a failed inspection system. It is very often a property of the sample rather than the appraisers.
This produces the single most common failure in attribute studies: parts pulled at random from normal production. Normal production is overwhelmingly conforming, which is the point of it, and a random sample therefore contains almost no defective units. Percent agreement comes out flattering, kappa comes out terrible, and neither number is informative.
The remedies are all about the sample rather than the statistic. Build the set deliberately with a substantial proportion of borderline and nonconforming units. Establish each unit's true state independently before the study, not by appraiser consensus afterwards. And report Po, Pe and kappa together, so a reader can see whether a low kappa reflects poor agreement or a lopsided sample.
Severity grades, cosmetic classes and 1–5 scales are ordinal: the categories have an order, and disagreements are not all equally serious. One appraiser calling a unit grade 2 while another calls it grade 3 is a minor disagreement. Grade 1 against grade 5 is not.
Ordinary kappa treats those identically, because it only asks whether the labels matched. That understates the performance of a system whose disagreements are all adjacent, and it hides the difference between a scale that needs slight tightening and one nobody understands.
Two better choices. Weighted kappa assigns partial credit for near-misses, with the weighting scheme stated in advance rather than chosen after seeing the results. Kendall's coefficient of concordance assesses whether appraisers rank units consistently, which is often the question that actually matters for a grading scale.
For a scale that fails, the diagnosis is nearly always in the pattern of disagreement rather than its amount: adjacent-category confusion points at boundary definitions, scattered disagreement points at a standard that is not operational.
Attribute studies need far more parts than variable studies, and the driver is the rare category, not the total. A measurement carries information in its value; a category carries only the fact of the category. That is a large loss, and sample size is where the bill arrives.
The arithmetic is unforgiving. At a 2% defect rate, fifty randomly chosen parts contain about one defective unit. Any statement about whether defects get detected then rests on a single observation. Tripling the sample to 150 parts yields about three — still nothing.
Sample deliberately instead. Construct a set with a substantial share of nonconforming and borderline units, confirm the true state of each independently, and present them blind and in random order across repeat rounds. That is not a random sample of production and should not be described as one — it is a designed test of the measurement system, and the protocol should say so.
The borderline units matter most. A study composed of obvious good parts and obvious scrap will pass comfortably and tell you nothing, because the decisions that go wrong in production are the ones near the boundary.
Whenever the underlying characteristic is continuous and can be measured, measure it. Converting a measurement into a pass/fail judgment discards most of the information the measurement contained.
The cost is not abstract. Variable data supports control charting, capability analysis and early warning of drift. Attribute data supports none of those — an attribute system reports nothing at all until the process crosses a threshold, at which point the problem is already producing scrap. It is also why attribute studies need so many more parts to say anything.
Attribute methods are the right tool where the characteristic is genuinely categorical: cosmetic defect class, presence or absence of a feature, a severity grade that no instrument produces. They are a fallback everywhere else, and a surprising amount of inspection is attribute by habit rather than necessity.
Where a measurement does exist, the questions become whether the gauge is adequate and what its noise does to your capability numbers.
You do not — a Gage R&R partitions variance, and a category has no variance to partition. The equivalent study is an attribute agreement analysis, which measures how often appraisers agree rather than how much they vary. It assesses three separate things: whether each appraiser agrees with themselves on repeat presentations, whether appraisers agree with each other, and whether they agree with a known reference standard. Those are different questions with different remedies, and a study that reports only overall agreement cannot distinguish them.
Because two appraisers who both call almost everything good will agree almost all the time by accident. If 98 percent of product passes, two inspectors who never look at the part and always say pass agree 98 percent of the time. Percent agreement does not subtract the agreement expected by chance, so on skewed product it reports a near-perfect measurement system that has demonstrated nothing. Kappa statistics exist to correct for chance agreement, which is why they are preferred — although they introduce a problem of their own on the same skewed data.
This is the kappa prevalence paradox, and it is usually a property of the sample rather than the inspectors. Kappa is observed agreement minus expected agreement, divided by one minus expected agreement. When one category dominates, expected agreement is already very high, so the denominator is small and even excellent observed agreement produces a low kappa. Two appraisers agreeing 95 percent of the time on product that is 97 percent conforming can post a kappa near zero. The fix is not to discard kappa but to fix the sample: deliberately include borderline and nonconforming units so the categories are better balanced, and report observed agreement, expected agreement and kappa together rather than kappa alone.
Far more than a variable study, and the number is driven by the rare category rather than the total. Attribute data carries much less information per observation than a measurement, and what the study actually needs is enough units in each category to estimate agreement within that category. If defects run at two percent and you take fifty random parts, you expect one defective unit and can conclude essentially nothing about whether defects are detected. Sample deliberately rather than randomly: build a set with a substantial share of borderline and nonconforming units, confirm each unit's true state independently, and present them blind and in random order.
Measure whenever measurement is possible. Converting a continuous characteristic into a category discards most of the information in it, which is why attribute studies need so many more parts and why attribute systems cannot detect drift until it crosses a threshold. A variable measurement supports control charting, capability analysis and early warning; a pass/fail judgment supports none of them. Attribute methods are the right tool when the characteristic is genuinely categorical — cosmetic defect class, presence or absence of a feature, severity grade — and a fallback everywhere else.