The clinical evaluation of software as a medical device (SaMD) follows the same legal basis as the clinical evaluation of other products, namely Article 61 and Annex XIV Part A of Regulation (EU) 2017/745, but requires its own evidentiary logic. The Medical Device Coordination Group described this logic in the MDCG 2020-1 guideline from March 2020. According to this, clinical evidence for software as a medical device consists of three components: valid clinical association, technical performance, and clinical performance. These components are not a step-by-step procedure, but rather methodological levels that together demonstrate that the software output is clinically sound, technically reliable, and effective within the healthcare context.
Valid Clinical Association
Valid clinical association, also referred to as scientific validity, describes the extent to which the software output is linked to the intended physiological state or clinical condition based on the selected input data and algorithms. The connection must be clinically recognized or scientifically well-founded and correspond to the state of the art. The guideline lists literature searches, technical standards, guidelines from professional societies, systematic reviews, proof-of-concept studies, clinical trials, and published data from summary reports, registries, and regulatory databases as data sources.
One example from the guideline is software that detects cardiac arrhythmias from digital auscultation signals: it must first be proven that the abnormal heart sounds used are actually linked to an arrhythmia. If this evidence is lacking, even excellent technical measurements will not support the clinical claim. For newly developed markers or data-driven models without an established literature base, the association must be proven independently, which regularly leads to the need for proprietary clinical data.
Technical and Clinical Performance
Technical performance is the ability of the software to reliably and correctly generate the intended technical output from the input data. According to MDCG 2020-1, evidence comes from verification and validation as part of good software development practice, from module, integration, and system tests, from tests in the intended computing and application environment, from curated databases and registries, and from previously collected patient data. The guideline lists performance characteristics including availability, confidentiality, integrity, reliability, accuracy (comprising trueness and precision), analytical sensitivity and specificity, limits of detection and quantitation, linearity, cutoff values, measuring range, generalizability, expected data rate and quality, the absence of unacceptable cybersecurity vulnerabilities, and human factors engineering.
Clinical performance is the ability to provide a clinically relevant output in accordance with the intended purpose. It may consist of a measurable, patient-relevant outcome, such as diagnosis, risk prediction, or prediction of treatment response, or a positive impact on patient management or public health. Validation must cover all intended purposes, target populations, conditions of use, operating and application environments, and all intended user groups. The guideline lists metrics such as clinical sensitivity and specificity, positive and negative predictive value, likelihood ratios, odds ratio, number needed to treat or harm, and confidence intervals. Validation at the module level is permissible if the module function is independent; new module combinations that change the indication or intended purpose require evaluation of the final configuration.
Distinction from Software Verification, Usability, and Follow-up
Software verification and validation according to the lifecycle model are prerequisites, but not a substitute for clinical evaluation. They show that the software meets its specification, not that the specification is clinically correct; this is only clarified by the clinical association. Similarly, the usability assessment answers whether users can operate the software with few errors, but not whether the output is clinically valid. MDCG 2020-1 links both levels technically by listing human factors engineering as a characteristic of technical performance, without naming individual software lifecycle or usability standards as evidence documents.
A further distinction must be made regarding follow-up: clinical evaluation is an ongoing process throughout the lifecycle and part of the quality management system; it is updated with data from the post-market clinical follow-up plan. For software, this feedback loop is particularly relevant because updates, changes to training data, new target populations, or altered operating environments can shift the previously demonstrated performance. Product categorization and classification itself is covered in the Medical Device Software entry.
Relevance for clinical trials
For software, this three-part division leads to a tiered study portfolio. Retrospective evaluations of curated datasets can support association and technical performance, while proof of clinical performance often requires prospective data from the intended healthcare setting, including actual user groups and real-world data quality. Central design questions include the definition of the reference, the independence of test data from development data, the pre-specification of thresholds, and the testing of generalizability across sites and device classes.
Because software is versioned, every data collection must be clearly assigned to a software version, and the follow-up plan must define which changes trigger a re-evaluation. Full-service CROs like Mediconomics support manufacturers in translating the three levels of evidence into concrete study and evaluation concepts, pre-defining reference standards and metrics, and documenting the link between software versions, study data, and follow-up in a traceable manner.
Frequently Asked Questions (FAQ)
Which three levels of evidence does MDCG 2020-1 require?
Valid clinical association (or scientific validity), technical performance, and clinical performance. Together, they constitute the clinical evidence for software as a medical device.
Is a retrospective dataset evaluation sufficient?
It can support association and technical performance, but usually does not cover clinical performance in the intended application environment and for all user groups. The required scope must be justified by the manufacturer in accordance with Article 61(1).
What triggers a re-evaluation after an update?
Changes affecting the intended purpose, target population, input data, algorithm, or application environment, as well as new module combinations that change the indication or intended purpose. In such cases, the final configuration must be evaluated.
Regulatory References
- Regulation (EU) 2017/745, Article 61 – Clinical evaluation, determination and justification of the required level of evidence.
- Regulation (EU) 2017/745, Annex XIV Part A – Clinical evaluation tasks including evaluation plan and data analysis.
- MDCG 2020-1, Guidance on Clinical Evaluation (MDR) / Performance Evaluation (IVDR) of Medical Device Software, March 2020.
- Regulation (EU) 2017/745, Annex VIII Chapter III Rule 11 – Classification of software as the basis for the scope of evidence.
- Regulation (EU) 2017/745, Annex XIV Part B – Post-market clinical follow-up plan.