Humans

Language model multimorbidity score predicts mortality in adults

Across eight external cohorts, the score yielded C-indices of 0.67 to 0.86, outperforming recalibrated conventional comorbidity indices by absolute gains of 0.12 to 0.21.

Fig. 1. Overview of the LLM-MMS: global validation, robustness assessment, and generation framework.
Open the figure at full size
Fig. 1. Overview of the LLM-MMS: global validation, robustness assessment, and generation framework.Overview of the LLM-MMS: global validation, robustness assessment, and generation framework.Liang et al.

Cyborg and Bionic Systems

In 770,439 adults from nine international longitudinal cohorts, researchers evaluated a large-language-model-based multimorbidity score (LLM-MMS) to predict all-cause mortality. The team converted harmonized health data into standardized natural-language narratives without model fine-tuning, using DeepSeek-V3 to estimate organ-specific biological ages, frailty age, and multidimensional health gradings. In the UK Biobank derivation cohort, the score achieved a C-index of 0.81, outperforming traditional comorbidity measures and raw-variable survival algorithms. Across eight external cohorts, LLM-MMS maintained discrimination with C-indices from 0.67 to 0.86 and observed-to-expected ratios between 0.94 and 1.04. Discrimination held across sexes and age groups, including adults under 60. Proteomic analyses linked higher score risk to elevated inflammatory and cellular-stress markers, such as GDF15, FGF21, and IL6.

Why it matters

The findings show that large language models can synthesize diverse routine clinical data into organ-specific biological ages and multimorbidity phenotypes without cohort-specific retraining. This approach captures systemic physiological decline and inflammatory pathways more effectively than static comorbidity indices across diverse populations.

Caveats

The analysis relied on observational cohorts with variable baseline disease definitions, missing data rates, and follow-up intervals ranging from 2.1 to 15.7 years. Furthermore, decision-curve analyses revealed attenuated or negative clinical net benefit at low threshold probabilities in two cohorts.

The paper

Bio-Inspired Artificial Intelligence Multimorbidity Score for Human-Machine Clinical Risk Stratification