Mediconomics – für individuelle CRO-Lösungen.

Glossar

The familywise error rate, often referred to as the Familywise Error Rate or FWER, is the probability of falsely rejecting at least one true null hypothesis within a predefined family of hypotheses. In confirmatory clinical trials, it is the benchmark ensuring that multiple efficacy claims do not arise from an accumulation of individual nominally significant tests. The family and the rule for its control are defined before the analysis.

Why multiple tests increase the error rate

With only one primary hypothesis, the significance level describes the risk of a false-positive conclusion for that specific test. If a programme includes multiple doses, endpoints, time points, or comparisons, there are multiple opportunities to obtain a favourable result by chance. The FWER considers this overall risk not on a test-by-test basis, but across the entire pre-specified set of conclusions.

The hypothesis family is a scientific and regulatory specification. For example, it may include claims for two doses and multiple endpoints, but it does not automatically include every exploratory analysis within a data package. The protocol and the statistical analysis plan must specify which decisions the family covers, the order in which testing is performed, and what conclusion a positive result supports in each case.

Methods to control the FWER

A simple allocation of the significance level across parallel hypotheses is only one possible approach. Hierarchical testing sequences, gatekeeping procedures, and closed testing procedures use the structure of clinical questions to release tests only when predefined requirements have been met. They are not error rates in their own right, but procedures whose property is control of the familywise error rate.

With co-primary endpoints, the clinical claim is often tied to the success of all specified tests. With multiple doses, an ordered sequence of dose comparisons can govern the transition to secondary objectives. Such rules must also take the planned confidence intervals into account so that estimation and the testing decision reflect the same multiplicity structure.

Control applies under all configurations in which one or more null hypotheses in the family are actually true. This distinguishes strong control of the FWER from protection that is considered only under the global null hypothesis. For clinical programmes, this distinction is practically relevant because a benefit at one dose or endpoint does not permit deriving an additional efficacy claim by chance when another null hypothesis is true. The chosen procedure must therefore fit the overall decision structure, not merely an expected pattern of results.

The FWER is therefore not a substitute for clinical prioritisation of endpoints. It does not impose a ranking, but it makes the formal implementation of that ranking verifiable. If multiple positive claims are made, it must remain clear which hypotheses actually belonged to the family and whether all predefined conditions for the respective claim were met. Documentation should also show which claims are explicitly exploratory only and therefore are not supported by the confirmatory testing procedure.

Distinction from single tests and basic concepts

The FWER is not the error probability of a single p-value. The p-value relates to the data under a specific null hypothesis; the familywise error rate assesses whether the set of planned testing decisions contains at least one false alarm. Also, null hypothesis (H0) and hypothesis describe the individual statement, not the error control of a set of statements.

A nominal p-value may be computationally small within a testing hierarchy and still not permit a confirmatory claim if the preceding hypothesis was not successful under the predefined rule. Adjustment is therefore not merely a post hoc correction step. It is part of the logic by which the study programme asserts an efficacy claim.

Relevance for clinical trials

Multiplicity affects study architecture as early as endpoint selection and sample size justification. If doses or subpopulations are later treated as additional confirmatory comparisons without a predefined rule, the interpretability of the original significance level is no longer assured. For interim analyses, repeated testing must likewise be included in error control.

Full-service CROs such as Mediconomics support the specification of the hypothesis family in the protocol, document testing hierarchies and decision rules in the statistical analysis plan, program the corresponding analyses, and verify that tables, listings, and the clinical study report present only those confirmatory claims covered by the multiplicity strategy.

Frequently Asked Questions (FAQ)

Does the FWER cover only multiple primary endpoints?

No. It may also apply to multiple doses, treatments, populations, or planned time points if confirmatory conclusions are to be drawn from them.

Are gatekeeping procedures and closed testing procedures synonyms?

No. Both can control the FWER, but they organise the transition between hypothesis families or the set of tests in different ways.

Does every exploratory analysis have to control the FWER?

An exploratory analysis may be descriptive or hypothesis-generating. If its result is intended to support a pre-specified confirmatory claim, its classification within the error-control framework must be clarified.

Regulatory References

  • EMA/CHMP/44762/2017, Guideline on multiplicity issues in clinical trials – explains the planning of multiple comparisons and the control of false-positive findings.
  • ICH E9, Statistical Principles for Clinical Trials – sets out hypothesis testing, error probability, and analysis in confirmatory studies.
  • ICH E20, Adaptive Designs for Clinical Trials – requires the maintenance of appropriate error control when adaptations are made.
Scroll to Top