medRxiv
Citation reliability of frontier large language models in medical writing and its automated verification
Preprint: computational studyAI & data
Abstract
Large language models (LLMs) are increasingly used to draft medical manuscripts, yet their citations are unreliable and clinicians lack a validated way to verify them. We evaluated three frontier LLMs, Claude Opus 4.8, GPT-5.5, and Gemini 3.
The paper
Seoul National University
medRxiv, 4 Sep 2026, CC BY, Preprint, not peer-reviewed