Ask a language model for medical references. Nearly half the time, it fabricates them entirely.
Gao, Zhang, Disis, and Zhang (arXiv:2603.22344) tested five LLM platforms on 2,000 references across 40 articles from major medical journals. The complete failure rate — returning references with no correct bibliographic data — was 47.8%. The average accuracy score was 0.29 out of 1.0. The best performer (Grok, 0.57) was roughly half-accurate. The worst (Gemini, 0.11) was noise.
The failure is not about getting a date wrong or misspelling an author. Complete failure means the reference doesn't exist — the title is invented, the authors are fictional, the journal is wrong. The model generates something that looks like a citation because it has learned the format, but the content is confabulated. The form is perfect. The substance is absent.
The variation across platforms (0.11 to 0.57) and journals (NEJM harder than BMJ) suggests the failure isn't uniform. Some platforms have better retrieval or grounding mechanisms. Some journal citation styles are easier to confabulate plausibly. The accuracy correlates with how well the training data covered that specific journal — which is itself a form of information: the model's confidence doesn't track its accuracy, it tracks its exposure.
For medical research, this is not a minor inconvenience. A fabricated citation that passes visual inspection — correct format, plausible title, real-sounding authors — can propagate through literature reviews, grant applications, and clinical decision support without detection. The failure mode is precisely that the output looks authoritative. The form carries the authority that the content lacks.