Mediconomics – für individuelle CRO-Lösungen.

Power Calculation

Statistical power is the probability that a statistical test detects an actually existing effect defined in the planning. For a given test, it corresponds to the probability of rejecting the null hypothesis when it is false, and is expressed as 1 minus beta. In clinical trials, power is a central parameter of the prospective sample size planning, but is not to be equated with the regulatory sample size justification.

Parameters of power planning

Power does not result from a single freely selectable number. Its calculation depends in particular on the clinically relevant effect size, the variability or event rate of the primary endpoint, the planned sample size, the significance level, the test method and the analysis strategy. For time-to-event endpoints, the expected number of events is often essential; for continuous endpoints, dispersion and measurement accuracy play a special role.

All assumptions must match the target population and the study objective. A variance or event rate taken from external data can be inappropriate if the population, standard therapy or data collection differ significantly. Likewise, the difference aimed for in the planning must be clinically and scientifically justified. ICH E9 requires the essential features of the statistical analysis to be described in advance in the protocol; therefore, for interpretation, not only the formula counts, but the comprehensibility of its inputs.

Types of error, endpoints and multiplicity

The type I error designates the risk of finding an effect although none exists under the null hypothesis. The type II error, beta, designates the risk of not detecting an existing effect with the intended test. Power is the counterpart to this second error. It does not describe a guarantee for the success of a single study and also does not state how large or clinically meaningful an observed effect is.

Multiple primary endpoints, multiple treatment comparisons or planned interim analyses influence the planning. If the overall probability of false positive conclusions is to be controlled, the testing strategy and possible adjustments must be defined in advance. This can change the required sample. Dropouts, non-evaluable data and an expected proportion of missing endpoints also belong in the planning, so that the target number of evaluable subjects can actually be achieved.

Differentiation from sample size justification and significance level

Statistical power is a property of the planned trial under certain assumptions. The sample size justification, on the other hand, is the regulatory and scientifically comprehensible presentation in the protocol of why the planned number of participants is appropriate. It documents, among other things, the endpoint, effect assumption, variance or event rate, power, significance level, test procedure, handling of multiplicity and assumptions regarding dropouts. The sample size justification therefore remains a separate glossary term; it is the documented evidence, not the power itself.

The significance level must also be distinguished from power. It defines in advance the accepted risk of a type I error for the test decision. Power, on the other hand, refers to the chance of detecting a defined effect. A change in the significance level can influence power with the same sample size, but does not replace a justification of the effect size or the targeted power. Both values must be planned consistently together with the hypothesis and the analysis method.

Relevance for clinical trials

An inadequately planned power can lead to a study not providing sufficiently precise or convincing evidence despite a relevant effect. An oversized study, on the other hand, can bind more participants than scientifically necessary. According to ICH E8(R1), appropriate study planning, participant protection and the focus on data essential for decision-making belong together. Under Regulation (EU) No 536/2014, the protocol must contain the statistical aspects including the justification for the number of subjects.

Full-service CROs such as Mediconomics support the definition of statistically robust assumptions, the selection of a procedure matching the endpoint, the programming and documentation of the sample size calculation as well as the protocol and statistical analysis plan. This includes sensitivity considerations for variance, event rates and dropouts, the review of multiplicity strategies as well as the coordination between clinical development, data management, biostatistics and medical writing.

Frequently Asked Questions (FAQ)

Does high power automatically mean clinical relevance?

No. Under the assumed conditions, high power increases the chance of detecting a defined effect. Whether the effect is clinically meaningful depends on the endpoint, magnitude, benefit-risk assessment and the previously justified objective.

Can power simply be increased after the start of the study?

A change in sample size or test strategy can influence the error probabilities and the integrity of the study. Such adjustments require pre-defined rules, a methodological justification and, if necessary, an amendment to the study documents; they must not be outcome-driven.

Why are historical data not always sufficient for planning?

Historical data can concern other populations, measurement methods, therapeutic environments or endpoint definitions. Their transferability must be checked and the uncertainty of the assumptions must be taken into account in the planning.

Regulatory references

  • ICH E9 “Statistical Principles for Clinical Trials” – sets out principles for the design, conduct, analysis and reporting of clinical trials.
  • ICH E8(R1) “General Considerations for Clinical Studies” – requires quality-oriented planning with a focus on factors relevant for decision-making.
  • ICH E6(R3) “Good Clinical Practice” – requires scientifically sound studies with reliable results.
  • Regulation (EU) No 536/2014 on clinical trials on medicinal products for human use – requires information on statistical aspects and on the justification for the number of subjects in the protocol.
Scroll to Top