Count-based records top foundation models for dementia risk
At 36 months before diagnosis, count-based models reached an AUROC of 0.738 compared with 0.719 for pretrained foundation models in All of Us records.
Research Square
In longitudinal electronic health records from human participants in the All of Us research program, researchers evaluated patient representation strategies for predicting Alzheimer’s disease and related dementias before diagnosis. The authors compared interpretable count-based models against pretrained clinical foundation models across several lead times and dementia definitions.
Count-based models achieved the strongest overall discrimination and calibration. Accuracy generally declined at longer horizons, where differences between representations narrowed. At 36 months before diagnosis, pretrained representations achieved an AUROC of 0.719 versus 0.738 for count-based models, along with higher sensitivity and F1 scores at a fixed operating threshold. Long-range predictions relied more on healthcare-utilization indicators, whereas cognitive phenotypes gained importance nearer diagnosis. In zero-shot testing on independent records from UChicago Medicine, performance degraded substantially across all models.
Why it matters
Identifying dementia risk years before clinical diagnosis could help target early interventions during early neurodegenerative changes. These findings indicate that simpler clinical count models can match or exceed complex foundation models while revealing the difficulty of transferring predictive tools between hospitals.
Caveats
The study relies on retrospective coding in electronic health records, which can be sparse or inconsistent, and models lost substantial accuracy when applied across health systems without retraining. This work is a preprint and has not yet completed peer review.
The paper
Institute for Population and Precision Health
Research Square · 1 Oct 2026 · Preprint, not peer-reviewed