medRxiv

Citation reliability of frontier large language models in medical writing and its automated verification

Preprint: computational studyAI & data

Abstract

Large language models (LLMs) are increasingly used to draft medical manuscripts, yet their citations are unreliable and clinicians lack a validated way to verify them. We evaluated three frontier LLMs, Claude Opus 4.8, GPT-5.5, and Gemini 3.

The paper