HumansPreprint

Proteomic disease prediction models generalize across biobanks

In models trained on 53,026 UK Biobank participants, 13 of 15 incident disease models maintained performance across FinnGen and All of Us cohorts.

medRxiv

In a preprint analyzing human data across three major cohorts, researchers tested how well blood proteomic disease risk models transfer across populations. The team trained prediction models for 15 diseases using Olink proteomics from 53,026 UK Biobank participants. Across the UK Biobank, the models reached a mean AUC of 0.74 (range 0.56 to 0.89), outperforming clinical factors alone by a mean Delta AUC of 0.03. Simple L2 models performed as well as complex transformer architectures. The researchers then evaluated the models in two external cohorts: FinnGen (n=5,865) and All of Us (n=7,405). Across biobanks, 10 of 14 prevalent models and 13 of 15 incident five-year disease models showed no significant drop in performance. Differences in phenotyping quality across cohorts were identified as the primary drivers of performance variation.

Why it matters

Proteomic profiles track dynamic physiological health states, and demonstrating their portability across distinct populations supports the development of scalable molecular tools for broad disease risk detection.

Caveats

This study is a preprint that has not yet been peer-reviewed, relies on observational biobank datasets, and notes that differences in phenotype definitions between healthcare systems can impact model accuracy.

The paper

Generalizability of proteomic risk prediction across biobanks reveals dependence on phenotype definitions