What the study measured
A team from UCSF, published August 4, 2026 in RSNA's Radiology, asked a narrow but consequential question: can a large language model write a better clinical indication for an imaging order than the referring clinician who placed it? They pulled 740,867 clinical notes across 77,626 imaging exams from 28,313 patients treated between January 2012 and August 2024, as reported by Radiology Business.
For each exam, the LLM received the original clinician-written indication plus the patient's ten most recent clinical notes, and generated an enhanced version. Twenty radiologists then blind-rated 25 indications apiece across four dimensions — comprehensiveness, factuality, conciseness, and usefulness — without knowing which version came from a clinician and which came from a model.
The gap was not close. Claude 3.5 Sonnet's indications were ranked the most comprehensive 37.14% of the time, versus 6.64% for the clinician-written originals, and the most factual 68.05% of the time versus 50% for clinicians. A smaller open model, Qwen 2.5-7B Instruct, still outscored clinicians on both counts — 28.42% top-comprehensiveness and 59.75% top-factuality. Comprehensiveness was also the single biggest driver of how radiologists ranked an indication overall, cited by 65.77% of readers as their primary factor.
Not a one-off finding
A similar pattern showed up a year earlier in a different setting. A 2025 study in Radiology from a large Canadian cancer center had GPT-4 draft oncologic clinical histories for 200 CT patients from existing chart notes. The AI-generated histories more consistently included the primary cancer diagnosis, treatment history, and acute or worsening symptoms than the original requisition text — and radiologists strongly preferred the AI versions for interpretation and safety.
Lead author Rajesh Bhayana, MD, put the stakes plainly: if a requisition for cancer imaging doesn't mention the primary tumor, "how we interpret the study and the likelihood we pick up cancer may be impacted." Two studies, two model generations, two very different clinical settings — and the same conclusion: referring clinicians, working under time pressure and often from memory, tend to under-document context that a model can pull from the chart in seconds.
Referring-clinician text vs. LLM-generated indication
| Metric (top rank) | Clinician indication | Claude 3.5 Sonnet | Qwen 2.5-7B |
|---|---|---|---|
| Comprehensiveness | 6.64% | 37.14% | 28.42% |
| Factuality | 50% | 68.05% | 59.75% |
Top-rank share among 20 blinded radiologist readers scoring 25 indications each. Source: Radiology (2026), as reported by Radiology Business.
The overlooked accuracy lever
Most discussion of AI reporting accuracy centers on the model reading the images: sensitivity for a specific finding, hallucination rate, agreement with a reference radiologist. That's the right thing to measure, but it treats the clinical indication as a fixed input rather than a variable one — and this research says it isn't fixed at all. It's inconsistent, often thin, and measurably improvable.
That matters because every reporting system — a radiologist, an AI drafting engine, or the two together — interprets a scan against the clinical question it was handed. An order that says "abdominal pain, rule out abnormality" gives a reader far less to work with than one that surfaces a prior surgery, a known malignancy, or a relevant lab trend buried three notes back in the chart. A more capable model reading against a thinner indication can still miss what a less capable model would have caught with better context. Order-quality and model-quality are two separate levers on the same output, and the field has mostly measured only one of them.
In practice, closing that gap doesn't require rewriting how referring physicians place orders. It requires a reporting pipeline that pulls the same kind of context an LLM pulled in this study — recent notes, prior imaging, known diagnoses — before the reading happens, rather than relying solely on whatever fits in an order's free-text field.
Where this fits with AI CT reporting
A foundation-model reporting system is exposed to the same order-context gap this study describes, which is why the drafting step should draw on all the clinical context available for a case rather than treat the order field as the only input, and why a review layer matters as a backstop when that context is incomplete. That's the structure behind AI CT reporting today: a comprehensive draft, an in-house radiologist review on every preliminary, and a ready-to-sign report for the client's reading radiologist. Better upstream context and a capable model aren't competing fixes for report quality — they compound.
Frequently asked questions
Do LLMs write better clinical indications than referring clinicians?
In a study published August 4, 2026 in Radiology, LLM-generated clinical indications for imaging orders were rated more comprehensive and more factual than the indications written by referring clinicians. Claude 3.5 Sonnet was ranked the most comprehensive indication 37.14% of the time versus 6.64% for the original clinician-written text, and the most factual 68.05% of the time versus 50% for clinicians.
What did the UCSF clinical-indication study measure?
Researchers analyzed 740,867 clinical notes across 77,626 imaging exams from 28,313 patients treated at UCSF between January 2012 and August 2024. LLMs were given the original clinician indication plus a patient's 10 most recent clinical notes and asked to generate an enhanced indication. Twenty radiologists then rated 25 indications each on comprehensiveness, factuality, conciseness, and usefulness, without knowing which indications were AI-generated.
Does this mean referring doctors are bad at writing imaging orders?
Not exactly — it reflects time pressure and workflow, not carelessness. Referring clinicians write indications in seconds during a busy visit, often from memory, while an LLM can pull from a decade of chart history in the same time. A similar pattern showed up in an earlier 2025 Radiology study, where GPT-4-generated oncologic histories more consistently included the primary cancer diagnosis, treatment history, and acute symptoms than the original requisition text.
Why does clinical indication quality matter for AI radiology report accuracy?
Any reporting system — human or AI — interprets images against the clinical question it's given. A thin or generic indication limits what a radiologist or an AI drafting engine can prioritize, no matter how capable the underlying model is. Closing the order-context gap is an accuracy lever that sits upstream of model quality, alongside detection performance and reviewer oversight.
Source: Serapio, Chen, et al., Radiology (2026), DOI: 10.1148/radiol.253238, as reported by Radiology Business. Corroborating study: Bhayana et al., Radiology (2025), DOI: 10.1148/radiol.242134. Figures are rounded as reported.