An analysis population is the predefined set of study participants and data used for a particular statistical analysis. It specifies who is included in an analysis and according to which rules data are included or excluded. Analysis populations must be appropriate for the study objective, the Estimand and the respective analysis; their criteria must be included in the protocol or statistical analysis plan and must not be changed based on results.
Populations according to analysis purpose
A clinical trial may require several analysis populations, for example for the primary efficacy analysis, for safety or for supplementary sensitivity analyses. Each population answers a comprehensible question only if inclusion criteria, assignment to treatment groups and the handling of missing data are clearly defined. ICH E6(R3) requires the selection of participants included in or excluded from planned analyses to be specified in advance and exclusions to be clearly justified and documented.
The definition alone is not sufficient. Protocol deviations, treatment initiation, endpoint availability and premature study discontinuation must be represented in a traceable manner in databases. Changes to analysis sets made after unblinding are justified only in exceptional circumstances and must be documented with particular transparency. A population tailored retrospectively may jeopardize the comparability of the groups and the credibility of the conclusion.
For the safety assessment, a population is often defined that comprises all persons who received at least one actual dose; however, the precise definition must be appropriate for the product and study conduct. Additional analysis sets may be required for laboratory values, exposure, patient-reported endpoints or pharmacokinetic data. Consistent designations are less important than fully documented criteria and clear assignment to the respective study objective.
Missing data are not a reason to silently remove participants from the primary analysis. The EMA points out that excluding incomplete cases may impair the comparability of the groups and bias estimates. For each relevant analysis set, it must therefore be described in advance how missing endpoint and covariate data will be handled and which sensitivity analyses will assess robustness.
Distinction from Intention-to-Treat, Per-Protocol and Full Analysis Set
Analysis population is the umbrella term. Intention-to-Treat refers to an analysis principle according to which the results of all randomized participants should be analyzed according to the randomized treatment group, irrespective of adherence or protocol compliance. It is therefore not simply an arbitrary list of persons, but a principle for preserving the randomized comparison.
Full Analysis Set is a specific designation used in ICH E9 for an analysis set that comes as close as possible to the Intention-to-Treat principle. Per-Protocol, by contrast, refers to a set with predefined and justified restrictions, for example excluding major protocol deviations. These terms are not interchangeable: which persons are included, how intercurrent events are taken into account and which Estimand is of interest must each be explicitly specified.
Relationship to Estimand and sensitivity analysis
According to ICH E9(R1), the Estimand describes the treatment effect intended to reflect the clinical question. The analysis population is only one component of the implementation: it does not alone determine how treatment discontinuations, subsequent therapies or death are addressed in the question. The selected estimator and the rules for missing data must be consistent with the Estimand.
In confirmatory trials, the primary analysis is often conducted in the Full Analysis Set, while Per-Protocol analyses or alternative assumptions can additionally assess robustness. Such analyses should not retrospectively select the most favorable finding, but should examine uncertainties defined in advance. Divergent results must be explained; they may limit the interpretation of a claimed treatment effect.
Relevance for clinical trials
Analysis populations link the operational conduct of a study and the statistical conclusion. Already during planning, it must be clear which data will continue to be collected after treatment discontinuation, how major protocol deviations will be assessed and how treatment codes will be transferred into the analysis. Regulatory authorities expect a complete derivation of the sets as well as consistent presentation in tables, listings and the clinical study report.
Full-service CROs such as Mediconomics support the definition of analysis sets in the protocol and statistical analysis plan, data validation and the reconciliation of randomization, treatment and endpoint data. They program traceable analyses, document inclusion and exclusion rules and coordinate sensitivity analyses and their description in the clinical study report.
Frequently Asked Questions (FAQ)
Is there only one analysis population per study?
No. A study may use different, prospectively justified sets for efficacy, safety and supplementary analyses.
May the primary analysis comprise only complete cases?
In confirmatory trials, this is generally not recommended as the primary analysis because it may violate the Intention-to-Treat principle and cause bias.
Is a Per-Protocol analysis always stricter?
It is more restricted, but not automatically more informative. Exclusions may alter the randomized comparison and must be justified in advance.
Regulatory references
- ICH E9, Statistical Principles for Clinical Trials — describes the Full Analysis Set and analysis principles in confirmatory trials.
- ICH E9(R1), Addendum on Estimands and Sensitivity Analysis in Clinical Trials — links the analysis to the prospectively defined Estimand.
- ICH E6(R3), Guideline for Good Clinical Practice — requires prospectively defined criteria for analysis sets and their documentation.
- EMA Guideline on Missing Data in Confirmatory Clinical Trials — explains the handling of missing data in the Full Analysis Set.