Global federated FAIR health data sharing infrastructures
Clinical Data Science performs research, develops and implements federated FAIR data infrastructures.
Federated data infrastructures mean that we focus on keeping the data at the source. This alleviates data privacy and control concerns.
FAIR data means Findable, Accessible, Interoperable and Reusable data. FAIR data is crucial in a federated data infrastructure as data may be captures in different languages and with different, local practices. Making health data available for learning in a federated FAIR manner, is one of our scientific strengths.
To further strengthen this theme, we also have extensive experience in dealing with legal, ethical and technical challenges including compliance to GDPR and related regulations and meeting security standards that our partners set.
Our approach is called the Personal Health Train:
Publications in this theme
Distributed learning on 20 000+ lung cancer patients - The Personal Health Train
Materials and methods: Lung cancer patient-specific databases (tumor staging and post-treatment survival information) of oncology departments were translated according to a FAIR data model and stored locally in a graph database. Software was installed locally to enable deployment of distributed machine learning algorithms via a central server. Algorithms (MATLAB, code and documentation publicly available) are patient privacy-preserving as only summary statistics and regression coefficients are exchanged with the central server. A logistic regression model to predict post-treatment two-year survival was trained and evaluated by receiver operating characteristic curves (ROC), root mean square prediction error (RMSE) and calibration plots.
Results: In 4 months, we connected databases with 23 203 patient cases across 8 healthcare institutes in 5 countries (Amsterdam, Cardiff, Maastricht, Manchester, Nijmegen, Rome, Rotterdam, Shanghai) using the PHT. Summary statistics were computed across databases. A distributed logistic regression model predicting post-treatment two-year survival was trained on 14 810 patients treated between 1978 and 2011 and validated on 8 393 patients treated between 2012 and 2015.
Conclusion: The PHT infrastructure demonstrably overcomes patient privacy barriers to healthcare data sharing and enables fast data analyses across multiple institutes from different countries with different regulatory regimens. This infrastructure promotes global evidence-based medicine while prioritizing patient privacy.
Keywords: Big data; Distributed learning; FAIR data; Federated learning; Lung cancer; Machine learning; Prediction modeling; Survival analysis.