Distilled language models improve multimorbidity scoring
Compact models trained on synthetic UK Biobank cohorts predicted patient survival with concordance indices reaching 0.91 without exposing real clinical data.
medRxiv
In health records from the UK Biobank, researchers developed a privacy-preserving computational framework to evaluate multimorbidity using large language models. The investigators distilled multimorbidity reasoning from large teacher models into compact student models, termed CoLLMs, using synthetic cohorts that replicated UK Biobank data distributions without exposing real patient records. The distillation achieved knowledge transfer with Spearman rank correlations ranging from 0.75 to 0.89. When evaluated on real UK Biobank data, CoLLM-derived multimorbidity scores improved survival prediction, yielding a concordance index of up to 0.91. The resulting scores also showed higher single-nucleotide polymorphism heritability, at an h² of approximately 0.05. An independent language-model evaluation verified the clinical significance of the distilled knowledge while identifying substantial performance variability across different teacher models.
Why it matters
Multimorbidity reflects the systemic accumulation of chronic conditions during biological aging. Privacy-compliant machine learning frameworks may allow researchers to capture complex disease interactions and predict mortality without risking the exposure of sensitive patient records.
Caveats
The study is a computational preprint that has not yet been peer-reviewed. The accuracy of the distilled student models depends heavily on teacher models, which showed substantial variability in their clinical reasoning.
The paper
Privacy-Preserving Distilled Large Language Models Enhance Multimorbidity Scoring
Case Western Reserve University
medRxiv · 21 Sep 2026 · Preprint, not peer-reviewed