A graphical testing procedure is a framework for controlling multiplicity when testing multiple hypotheses in a clinical trial. Nodes represent elementary null hypotheses, while initial weights assign local significance levels, and directed, weighted edges regulate the propagation of this level after a rejection. The graph thus visually represents the pre-specified testing strategy, rather than describing it solely through a long series of case distinctions.
The Graph as Test Architecture
Each hypothesis initially receives a portion of the available local significance level. This may concern, for example, the primary endpoint, a secondary primary endpoint, and selected secondary endpoints. A directed edge specifies which proportion of the released level is passed on to other hypotheses if the associated node has been successfully tested and rejected.
The weights and edges express a scientific hierarchy: a key endpoint can be tested first, a confirmatory endpoint only receives a level after its success, or a proportion is distributed among several questions. The graph is therefore not merely a decorative illustration. Its numerical values determine which claims are confirmatory and permissible for which result pattern, and must align with study objectives and endpoint priorities.
Iterative Process and Documentation
At each step, it is checked whether the p-value of a currently testable hypothesis reaches its current local level. Upon rejection, its weight is redistributed according to the outgoing edges; hypotheses without a received level remain untestable in this step. The process ends when no further rejection meets the specified conditions. The overall logic thus prevents residual levels from being used multiple times.
In the study protocol, the graphical representation should be supplemented by an unambiguous tabular specification. This includes node designations, initial weights, edge directions, edge proportions, testing order, and rules for all combinations of relevant decisions. For reproducibility, the programming requires a machine-readable parameter table; a graphic drawn alone leaves ambiguities regarding rounding, nodes without edges, or special rules for missing evaluability.
If a node has multiple outgoing edges, the proportions only become active after this node has been rejected. If the test fails, the local level bound to it remains there and cannot be passed on to subsequent hypotheses. This asymmetry distinguishes a hierarchical graph from a retrospective ranking of favorable p-values.
For co-primary endpoints, the graph must also reflect the logical success rule. If both hypotheses are to be successful, a mere distribution of the level to two independent nodes is insufficient; the specification must show how the joint requirement is tested. The graph therefore does not replace a scientific decision on whether endpoints are alternatively, hierarchically, or jointly critical for success.
Distinction: Not a Single Test
A graphical testing procedure is not a single statistical test, such as a t-test or a log-rank test. It controls how local levels are applied across multiple such individual tests. The p-values used can originate from different models, each appropriate for the endpoint, as long as the multiple testing procedure defined in the graph is observed.
Gatekeeping, fixed sequences, and closed testing procedures are not opposing concepts, but can be represented as special cases or closely related sequential-rejective procedures, depending on their structure. The advantage of the graph lies in the transparency of complex dependencies: instead of listing every intersection hypothesis in prose, it shows which rejection activates further tests and where the local level flows.
Relevance for clinical trials
Multiple dosages, co-primary endpoints, or hierarchical secondary objectives require a testing strategy that is fixed before data unblinding. Clinical development, biometrics, and regulatory affairs must therefore clarify which endpoint should serve as proof of efficacy and which claims may only be tested subsequently. Changes to edges or weights after knowledge of results alter the confirmatory interpretation and are not merely a format correction.
For interim analyses, it must also be determined whether and how an alpha expenditure over time points is combined with the distribution across hypotheses. A multiplicity graph regulates the second dimension but does not, without explicit extension, adopt the rules of a group-sequential design.
Full-service CROs like Mediconomics support graphical testing procedures by translating the endpoint strategy into a formal graph specification, aligning the rules with the SAP, independently validating the statistical programs, and tabulating the test paths in the CSR. This allows for tracing the local threshold available at the time for each tested hypothesis upon database lock.
Frequently Asked Questions (FAQ)
What does a weight of zero at a node mean?
The hypothesis is not confirmatorily testable at the outset. It can only receive a positive local level if one or more upstream hypotheses have been rejected and the graph’s edges provide for propagation to this node.
Can multiple hypotheses receive a local level simultaneously?
Yes. A graph can distribute the initial level or proportions released after a rejection to multiple nodes. Especially then, the sum of the propagation weights and the rule for each subsequent test must be precisely specified.
Does the graphic replace the written multiplicity section in the SAP?
No. The graphic facilitates scientific review, but the SAP must define the hypotheses, p-values, weights, edges, transition rules, and analysis set in such a way that implementation and verification are unambiguously possible.
Regulatory References
- EMA/CHMP/44762/2017, Guideline on multiplicity issues in clinical trials – addresses the control of false positive conclusions.
- ICH E9, Statistical Principles for Clinical Trials – establishes the pre-specification of statistical analysis.
- ICH E9(R1), Estimands and Sensitivity Analysis – links study objective and precise analysis question.