A statistical analysis plan that survives review

An SAP is not a description of the analysis. It is a commitment made before the data are seen — and almost everything that makes one hold up or fall over is decided by what it commits to, and when.

Short answer

The document has one job: to make it verifiable that the analysis reported was the analysis intended. Everything else follows from that.

Which means the binding constraint is timing, not content. A plan finalized before database lock and before anyone with influence over the analysis has seen unblinded data carries evidentiary weight. The identical document written afterwards does not, however good it is, because it documents choices that could have been informed by the results.

Version, date and approval should be unambiguous in the record. That is not administrative tidiness — it is the entire basis of the claim.

The estimand comes before the method

The most common structural weakness is a plan that specifies a model without specifying the question. The same endpoint supports several different treatment effects, and which one is being estimated is a scientific decision that the analysis method should follow from rather than imply.

A complete estimand states the population, the endpoint, how intercurrent events are handled, and the summary measure. The intercurrent-event component is the one most often left implicit and the one that changes the answer most: whether a subject who discontinues treatment, or takes rescue medication, contributes their observed outcome, a hypothetical outcome, or nothing at all defines a different effect each time.

Getting this explicit early has a practical benefit beyond compliance. It surfaces disagreement between the clinical and statistical teams while it is still cheap — disagreement that otherwise emerges at the analysis stage, when it is not.

Populations, and what they are for

Define analysis populations by rule rather than by judgment, and state which analyses run on which. Intention-to-treat preserves randomisation and is conservative for superiority; per-protocol is not protected by randomisation and can bias in either direction, which matters especially in a non-inferiority setting where both are usually required.

The rule that decides population membership must be applicable without knowing the outcome. A criterion that requires looking at a subject's result to determine whether they belong is not a population definition — it is a selection mechanism, and it will be treated as one.

Missing data: state the assumption

This section receives disproportionate scrutiny because it is where plans most often assume the convenient thing quietly.

Complete-case analysis assumes data are missing completely at random. That is rarely plausible when subjects discontinue for reasons connected to the outcome, which is the usual situation. Choosing an approach without naming the assumption leaves the plan resting on something it never claimed.

The defensible structure has three parts: a stated assumption, a primary analysis consistent with it and with the estimand, and sensitivity analyses pre-specified in advance that probe what happens if the assumption fails. Sensitivity analyses selected after seeing how much data went missing, and which direction it went, are not sensitivity analyses.

Multiplicity, decided in advance

Decide before the study which claims require type I error control, and specify the procedure completely — including the ordering, because the ordering is where the judgment lives.

A hierarchical strategy fixes a sequence and stops at the first failure. It costs nothing when the ordering reflects genuine clinical priority and a great deal when the ordering was chosen hopefully. Alpha splitting and gatekeeping procedures trade differently. Any of them is defensible; what is not is deciding afterwards which endpoints were confirmatory, and that is usually visible.

The same discipline applies to interim analyses. If the study may stop early, the spending function, the timing, and who sees what all belong in the plan — along with the firewall between the people who see interim data and the people who could act on it.

Amendments without damaging the record

Amending is ordinary. Amending badly is not.

An amendment that tightens an ambiguous definition before any unblinded data exists is unremarkable and improves the plan. An amendment after interim results, or one that changes the primary analysis without a reason independent of the data, raises a question that is difficult to answer well no matter how good the reason actually was.

The test worth applying: does this change what will actually be done? If yes, amend and record the reason and the date. If it merely restates reasoning already in place, do not — it creates version history someone will later ask about, for no benefit. That principle also governs the response when a regulator asks a statistical question, which is the subject of the note on answering an FDA Information Request.

What reviewers actually challenge

Most of these are cheap to prevent and expensive to fix, which is the argument for a statistical review of the plan before it is signed rather than a defence of it afterwards.

Send the problem, not the data →

Common questions

When does a statistical analysis plan have to be finalized?

Before database lock and before anyone with influence over the analysis has seen unblinded data. That is the whole basis of its evidentiary value: a plan written afterwards documents choices that may have been informed by the results, and no amount of quality in the document recovers what the timing gave away. The version, date and approval should be unambiguous in the record. A plan finalized late is not automatically fatal, but it converts a settled question into one the sponsor has to argue.

What is an estimand and why does it come before the method?

An estimand is a precise statement of the treatment effect being estimated: the population, the endpoint, how intercurrent events such as discontinuation or rescue medication are handled, and the summary measure. It comes first because the same endpoint can support several different questions, and the analysis method is a consequence of which question is being asked rather than a choice made independently. Plans that specify a model without specifying the estimand leave the actual question implicit, and reviewers increasingly ask for it directly rather than inferring it.

How should missing data be handled in a statistical analysis plan?

By stating the assumption, choosing a primary approach consistent with the estimand, and pre-specifying sensitivity analyses that probe departures from that assumption. Complete-case analysis assumes data are missing completely at random, which is rarely plausible when subjects discontinue for reasons related to the outcome. The defensible structure is a stated primary assumption, an analysis appropriate to it, and sensitivity analyses selected in advance rather than after seeing how much data went missing. Reviewers examine this section closely because it is where a plan most often turns out to have assumed the convenient thing.

How do you handle multiplicity across endpoints?

By deciding in advance which claims require type I error control and specifying the procedure completely, including the ordering. A hierarchical testing strategy fixes a sequence and stops at the first failure, which costs nothing when the ordering reflects genuine priority and costs a great deal when it was chosen hopefully. Splitting alpha across endpoints, or a gatekeeping procedure, are alternatives with different consequences. What is not acceptable is deciding after the fact which endpoints were confirmatory, and a reviewer can usually tell.

Can a statistical analysis plan be amended?

Yes, and amending before database lock is ordinary practice when an analysis proves underspecified. What matters is the record: what changed, when, why, and what was known at the time. An amendment that tightens an ambiguous definition before any unblinded data exists is unremarkable. An amendment made after interim results, or one that changes the primary analysis without a stated reason independent of the data, invites a question that is difficult to answer well. Amending to restate reasoning already in place creates version history someone will later ask about for no benefit.