Application of FAIR principles to the validation and interoperability of clinical AI models, using model-based selection for proton therapy as a case study
Artificial Intelligence (AI) models are increasingly being developed and deployed in healthcare, where they hold great potential to enhance clinical decision-making, improve diagnostic accuracy, and optimize patient outcomes. However, the integration of AI into clinical workflows requires transparency, reproducibility, and interoperability to ensure that these systems operate safely and effectively across different healthcare settings.
To address similar challenges in data management and processing, the FAIR (Findable, Accessible, Interoperable, and Reusable) principles were introduced as a framework to standardize the way scientific data is stored, shared, and reused. Initially designed for data, the FAIR principles have since been recognized as valuable for describing and managing AI models as well. Recent initiatives have explored the application of the FAIR principles to AI research data and AI models themselves. Moreover, a metric model has been proposed to quantitatively evaluate the level of FAIRness of AI models.
Previously, our group has developed the FAIVOR tool—a robust, privacy-preserving platform designed to evaluate and adapt pre-trained medical machine learning models to new datasets within hospital environments. FAIVOR enables models to be converted into a containerized format, ensuring interoperability across systems regardless of the original programming language or framework. This approach allows healthcare institutions to validate in a privacy-preserving manner.
According to the European AI Act and related international guidelines, AI models implemented in clinical practice must undergo periodic validation to ensure continued reliability, safety, and fairness. Building upon these requirements, this study proposes to adapt FAIR principles for specific, widely used AI models and subsequently to evaluate their FAIR metrics.
As a use case, we propose to work with Normal Tissue Complication Probability (NTCP) models. These models predict the likelihood of radiation-induced toxicities and are critical for optimizing radiotherapy (RT) treatment planning. Specifically, NTCP models are used to compare the expected normal tissue complication probability profiles between photon and proton RT treatment plans. Such comparisons allow clinicians to identify patients who are likely to benefit most from proton therapy (PT) in terms of reduced toxicity rates. The clinical decision is guided by the difference between proton and photon NTCP profiles, referred to as ΔNTCP, as described in the National Indication Protocol for Proton Therapy in the Netherlands (Version 2.2).
Given the significant impact of these models on patient outcomes and clinical decision-making, it is essential to verify their performance and generalizability on new datasets. Previous independent validations, such as the work by Kalendralis et al., have been limited to specific side effects like dysphagia. Comprehensive validation across additional clinical endpoints remains an unmet need.
Therefore, this study aims to validate a digital NTCP model developed in accordance with FAIR principles using a new dataset.
Research question:
1. Can we make the NTCP models for ProTRAIT FAIR?
2. What are the validation results of these FAIR models, based on retrospective data of Maastro Clinic?
Research design:
Analytical methods are applied to define FAIR metrics for research software for this model. This study employs a cross-sectional design for the validation of the models on retrospective data.
Methods for data collection:
This study is based on retrospective data of Maastro Clinic, collected for previous projects.
Potential supervisor:
Ekaterina Akhmad (PhD student) ekaterina.akhmad@maastrichtuniversity.nl
Dr. Johan van Soest (Assistant Professor) j.vansoest@maastrichtuniversity.nl
Required skills:
English language - required
Dutch language - Might be useful as some documents are in Dutch, but it’s not an obligatory requirement
Statistical analysis - Python, R using methods of descriptive statistics, tests for comparing distributions, and methods to build and analysis performance metrics of the model (Brier Scores, Calibration plots: graphical and quantitative assessments, discrimination evaluation: sensitivity, specificity, AUROC)
Programming skills - Preferred knowledge of Python, R
Qualitative data analysis - Not required
Time period: 3-6 months
Location: Maastricht, the Netherlands or remote