Method notes for regulators and policy analysts: what utilization management data can support, what it cannot, and how to structure an analysis that survives scrutiny from both directions.
Denial rates are not comparable across plans without adjustment. A raw rate is a ratio whose numerator and denominator both depend on plan behavior: a plan requiring prior authorization for more services generates more decisions and therefore more denials, without necessarily denying more care.
Any defensible comparison holds the requested service, the enrollee population and the request volume constant — service-line stratification and risk adjustment before any cross-plan statement is made. And the deepest limitation is not statistical: claims data cannot observe care that was never requested.
Claims data records requests that were made. It has no field for the request a physician's office decided not to submit because prior authorization for that service was known to be onerous, or because a previous attempt had failed.
That deterrence effect is invisible in every denial statistic ever computed. It does not appear as a denial, an appeal, or an unmet need — it appears as nothing at all, or as a substituted service that was easier to authorize. An analysis that presents the denial rate as a measure of access to care understates the burden by an unknown and unestimable amount.
The honest treatment is to state this plainly in the limitations and to resist the temptation to model around it. A sensitivity analysis cannot bound a quantity the data never observed. Saying so is more credible than a confidence interval that implies otherwise.
The denominator is a policy choice. Prior authorization requirements differ by plan and by service line. Two plans with identical clinical behavior will report different denial rates purely because one submits more decisions to the process.
Case mix differs. Enrollee populations vary in age, comorbidity and the mix of services they need. A plan enrolling a sicker population faces more requests for high-cost services, which are precisely the services subject to review.
Denial is not one event. An initial adverse determination, a partial approval, an approval with modified quantity, a denial later overturned on appeal and a denial for administrative incompleteness are all frequently pooled into "denials," and they mean different things. A pooled rate is not interpretable without knowing the mix.
Service lines are not exchangeable. Post-acute care, advanced imaging, durable medical equipment and specialty drugs have different clinical bases for review and different baseline approval rates. Averaging across them produces a number describing no actual decision.
Overturn rates on appeal are the most-quoted and most-misread statistic in this area.
A high overturn rate is genuine evidence about the quality of initial decision-making — if a large share of appealed denials are reversed, initial determinations are frequently wrong in a specific direction. That is a defensible finding.
What it does not support is an inference about all denials. Only a small fraction are appealed, and appealed cases are emphatically not a random sample: appeals concentrate where the clinical stakes are high, where the provider has administrative capacity, and where the case is strong. The overturn rate among appealed cases is therefore an upper bound on what the rate would be among all denials, not an estimate of it.
Presenting it as the latter is the most common error in this literature, and it is the one an opposing analyst will identify first.
Pre-specify the protocol. Write the hypotheses, the service-line definitions, the adjustment set and the comparison structure before touching the data. In a regulatory context this is not methodological hygiene — it is what allows a finding to withstand the claim that the analyst searched for a result.
Define the unit of analysis explicitly. The request, the enrollee, the episode and the claim line are four different units producing four different answers. Most disagreement about denial rates dissolves once both parties state which one they mean.
Stratify before you adjust. Regression adjustment across incomparable service lines produces a coefficient with no clinical referent. Stratify first, adjust within stratum, and report by stratum — a pooled estimate can follow, but it should not lead.
Report the denominator alongside every rate. A denial rate without the request volume it was computed from is not an interpretable quantity, and it is the first thing a competent reviewer will ask for.
MRP Group served as statistician on a study of the utilization management and provider payment practices of Medicare Advantage plans for the State of Connecticut, as a subcontractor to Evidence Impact Labs. The work covered authorship of the statistical protocol and hypotheses; preparation of CMS and Connecticut All-Payer Claims Database data; descriptive, comparative and regression analysis of utilization, cost and utilization-management impact; and analysis of claims denials, appeals and prior authorizations. Data integrity and intellectual property were managed in a controlled AWS S3 environment.
The report was published on the Connecticut Insurance Department portal in December 2024, with the full Statistical Analysis Plan as Appendix A1.
Not without adjustment. Both the numerator and the denominator depend on plan policy about which services require review. Service-line stratification and risk adjustment come before any cross-plan claim.
Care that was never requested. The deterrence effect of an onerous prior authorization process appears nowhere in the data, and no sensitivity analysis can bound it. It belongs in the limitations, stated plainly.
It is proof that initial determinations are frequently reversed, which is meaningful on its own. It is not an estimate of the rate among all denials, because appealed cases are a selected sample — it is an upper bound.
Typically CMS data for plan characteristics and enrollment, plus a state all-payer claims database for service-level utilization and payment. Neither alone supports plan comparison. Both usually require a controlled environment under state data-use agreements.
Narrower than instinct suggests. 'Are Medicare Advantage plans denying too much care' is not answerable from claims data. 'For this service line, among comparable enrollees, do plans differ in the proportion of requests denied at initial determination, and how much of that difference persists after adjustment' is answerable, and the answer supports a defensible finding. The value of the pre-specified statistical protocol is that it forces this narrowing before the data are seen.
Working on a utilization management or claims question?
A short call will establish what the available data can actually support. Get in touch →