In benefit assessments, certainty of conclusions refers to how strongly a statement about an added benefit or harm is scientifically substantiated. IQWiG distinguishes the levels indication, hint and proof. The scale does not describe the size of an effect, but how robust the conclusion is based on the available evidence.
The three levels of certainty
A proof is the strongest level, a hint an intermediate level, and an indication the weakest of the three positive levels of certainty of conclusions. The starting point is study quality, design, consistency of results, and whether relevant risks of bias limit interpretation. The label is assigned to a specific statement about benefit or harm; it is not a seal of quality for a medicinal product as a whole.
The scale also allows for an effect not to be proven and therefore to remain unclear. IQWiG’s methods aim not to extend the evidence beyond what the studies can support. For patient-relevant endpoints, it is therefore assessed whether results actually show differences in survival, symptoms, functioning, quality of life or adverse effects, and whether their measurement is credible.
Overall assessment of the evidence
Certainty of conclusions arises from an overall assessment, not from a single numerical result. Randomised studies can, for example, lose evidential value due to lack of blinding, high dropout rates, unsuitable analysis populations or deviating endpoint definitions. Several studies can support a conclusion provided their results point in the same direction and the studies are generalisable to the assessment question.
For a benefit assessment, certainty is considered for the respective endpoints and patient groups. This results in a substantiated overall assessment. A large observed difference does not automatically lead to a proof if the evidence base is uncertain. Conversely, a methodologically robust study can provide proof of a small difference.
Distinction: effect size and statistical metrics
A proof of a minor added benefit is not the same as an indication of a major added benefit. The former describes high certainty with a small magnitude; the latter describes weaker certainty with a potentially larger magnitude. The dimensions of probability and magnitude are assessed separately and only then combined into the overall conclusion.
Certainty of conclusions is also not a synonym for confidence interval, p-value or bias. These terms can be part of a study’s assessment: confidence intervals show the precision of estimates, p-values relate to a statistical test decision, and bias describes systematic distortion. Indication, hint and proof, by contrast, summarise the suitability of the overall evidence for a benefit claim.
The classification is endpoint-specific. For mortality, for example, there may be a hint, while for health-related quality of life only an indication—or no robust conclusion at all—may be possible due to incomplete questionnaires. The subsequent overall assessment must make these differences visible rather than applying a single evidence grade across an entire study programme. Where benefit and harm signals compete, not only certainty but also the patient-relevant weighting of the results is assessed.
The terms must therefore not be confused with the evidence levels of individual guidelines or with the quality of a publication. An article may be well reported and still not allow a conclusion on added benefit in Germany if the population studied or the comparator does not match the assessment question. Conversely, a clearly documented study limitation remains relevant to certainty of conclusions even if the results table appears statistically striking.
An explicit presentation of the reasons for uncertainty also makes the assessment useful for further evidence planning. It shows whether a design issue, a data gap or limited generalisability would need to be addressed.
The scale therefore protects against overinterpreting isolated results. It forces the strength of a conclusion to be tied to a robust evidence base and to classify positive and negative effects with the same methodological rigour.
Relevance for clinical trials
For study programmes, the scale means that it is not sufficient to focus solely on producing a statistically significant result. Protocol adherence, complete capture of patient-relevant endpoints, pre-specified analyses and an appropriate comparator therapy also determine how the data will later be classified. Different levels of certainty of conclusions in subgroups can arise when event counts, study coverage or generalisability diverge.
Full-service CROs such as Mediconomics support this through precise protocols, monitoring of data quality and documentation of pre-specified analyses. Biostatistics checks consistency and sensitivities, while Medical Writing explains the evidence base by endpoint; this allows transparent presentation of why a statement may reach an indication, hint or proof.
Frequently asked questions (FAQ)
Is a proof always associated with a major added benefit?
No. Proof describes the certainty of the conclusion, whereas major, considerable or minor relate to the magnitude of an added benefit.
Can a statistically significant endpoint yield only an indication?
Yes. Study limitations, limited generalisability or uncertainties in other aspects of the evidence can weaken the overall conclusion.
Is certainty of conclusions assigned equally for all patients?
Not necessarily. For separate patient groups, the available evidence may be robust to different degrees.
Regulatory references
- IQWiG, General Methods Version 8.0 – explains how levels of certainty of conclusions are derived.
- Section 35a SGB V – provides the legal framework for the added-benefit assessment.
- AM-NutzenV – contains categories for benefit and added benefit.
- Rules of Procedure of the G-BA, Chapter 5 – governs the assessment in the German procedure.