Scientific Data

PRIME-CVD: Paired Synthetic Cohort and Structured EMR-Style Data Assets for Education in Cardiovascular Risk Modelling

Figure 1 from Scientific Data
Open the figure at full size
Figure 1. Kuo et al., CC BY
Computational studyAI & data

Abstract

Abstract In recent years, progress in medical informatics and machine learning has been accelerated by the availability of openly accessible benchmark datasets. However, patient-level electronic medical record (EMR) data are rarely available for teaching or methodological development due to privacy, governance, and re-identification risks. This has limited opportunities for students from healthcare and clinical backgrounds to gain hands-on experience in coding, data management, and cardiovascular risk modelling. Here we introduce PRIME-CVD, a parametrically rendered informatics medical environment designed explicitly for medical education. PRIME-CVD comprises two openly accessible synthetic data assets, each representing 50,000 adults undergoing primary prevention for cardiovascular disease and parameterised to reproduce the statistical characteristics of a derivation cohort from the Australian New South Wales Lumos linked-data resource. The datasets are generated entirely from a user-specified causal directed acyclic graph parameterised using characteristics of the reference cohort, publicly available Australian population statistics and published epidemiologic effect estimates, rather than from patient-level EMR data or trained generative models. Data Asset 1 provides a clean, analysis-ready cohort suitable for teaching exploratory analysis, stratification, statistical coding, and survival modelling, while Data Asset 2 systematically transforms the same cohort into a structured relational, EMR-style database with deliberately introduced data-quality challenges, including missingness, heterogeneous terminology, inconsistent measurement units, and cross-table linkage requirements. Together, these complementary assets enable healthcare and clinical learners to develop practical skills in coding, data cleaning, harmonisation, cohort reconstruction, and cardiovascular risk modelling without exposing sensitive information. Because all individuals and events are generated de novo, PRIME-CVD preserves realistic subgroup imbalance and risk gradients while ensuring negligible disclosure risk. Although the released PRIME-CVD cohort represents adults with no prior history of cardiovascular disease undergoing primary-prevention risk assessment, the complete data-assembly pipeline is provided through the dedicated GitHub repository. Educators may adapt the supplied code, parameters, and data-generating assumptions to construct alternative synthetic cohorts suited to their own teaching objectives. PRIME-CVD is released under a Creative Commons Attribution 4.0 licence to support reproducible and scalable health data science education.