Method notes

Written answers to questions that come up repeatedly in regulatory and trial work, and that have thin coverage elsewhere. Each is drawn from an engagement rather than assembled from the literature.

Manufacturing · measurement systems

Why a Gage R&R gets rejected

Three quite different problems produce the same red number: parts that did not span what the process actually makes, the wrong acceptance ratio applied to the application, or a gauge that genuinely cannot resolve the tolerance. They call for opposite responses, and telling them apart rarely needs new data.

Clinical trials · pre-specification

A statistical analysis plan that survives review

An SAP is a commitment made before the data are seen, not a description of the analysis. The estimand comes before the method, the missing-data assumption has to be named, and timing carries more weight than polish.

Clinical trials · FDA

Answering an FDA Information Request on a vaccine IND

Most statistical IRs ask a sponsor to make an existing decision auditable, not to change it. The right response answers the question asked, in the reviewer's frame, with the smallest sufficient addition to the record — because a reply that generates a follow-up has cost a full review cycle.

Radiological survey design · NRC

Scan MDC for continuously collected data

Classical scan MDC assumes independent counts in discrete intervals. A continuous detector stream is serially correlated, so the effective number of independent observations is smaller than the raw count and the nominal false-positive rate is not the actual one. Generalized midpoint, moving-average and EWMA lag-k differencing restore control; bootstrap gives the bias and precision of the resulting limits.

Manufacturing · capability

Process capability through a noisy measurement system

Variances add, so observed spread always exceeds true process spread and Cpk computed from measured data understates a capable process. The correction is simple arithmetic; knowing which number the decision actually needs is not.

Medical devices · validation

Device agreement and test–retest reliability

A new instrument is almost never compared against truth — it is compared against another instrument that also carries error. Correlation answers the wrong question; bias, limits of agreement and reliability answer three different ones.

Veterinary biologics

Prevented fraction in veterinary vaccine efficacy

At challenge-study sample sizes the confidence interval matters more than the point estimate, because the lower bound is what supports the label claim. And the case definition does more work than any model — which is why it has to be pre-specified.

Measurement · cross-domain

What an accuracy percentage means

Ninety-eight percent accurate — against what, and measuring bias or spread? Comparing against an imperfect reference biases the answer in a knowable direction, and the tails, not the average, are where the argument happens.

Regulatory · conformity assessment

Setting a guard band

A specification limit stops being a decision rule once measurement carries error. What replaces it is a deliberate split of risk between producer and consumer — and presenting it as paired risk curves survives challenge far better than a single threshold.

Manufacturing · discrete and categorical

Attribute agreement analysis

A category has no variance to partition, so Gage R&R does not apply. Agreement does — within appraiser, between appraisers, and against a known standard. Percent agreement flatters and kappa collapses, both for the same reason: the sample.

Health policy · claims analysis

Medicare Advantage prior authorization and claims denials

Denial rates are not comparable across plans without adjustment, because the denominator is itself a policy choice. And the deepest limitation is not statistical: claims data cannot observe care that was never requested, so every denial statistic understates burden by an unknown amount.

Manufacturing · destructive testing

Nested Gage R&R when you cannot measure the same part twice

Destructive and one-shot testing break the standard crossed study, because the unit is consumed. The nested design substitutes homogeneity for repetition — and that substitution, not the arithmetic, is where the study can quietly go wrong.