← BlogAI Safety & OversightJuly 24, 20267 min read

    Could an LLM catch report errors
    your quality team misses?

    A 735-patient study on breast ultrasound reporting found a large language model beat manual reviewers on quality-control accuracy — and finished the job about 16 times faster. Here's what the numbers show, and where a QA layer like this fits alongside radiologist-reviewed reporting.

    735
    Patients studied
    across 60 hospitals in China
    80% vs 64%
    Margin accuracy
    LLM vs manual QC
    74% vs 56%
    Echo pattern accuracy
    LLM vs manual QC
    13 vs 213 min
    Time for 50 reports
    LLM vs manual review

    What the study tested

    Quality control (QC) — often grouped under the broader banner of radiology quality assurance — is one of the least glamorous, most time-consuming jobs in a radiology department: someone has to check that finished reports actually conform to standards like BI-RADS, catching errors, inconsistencies, and omissions before they cause downstream problems. A team led by corresponding author Hongyan Wang, of the department of ultrasound at Peking Union Medical College Hospital in Beijing, tested whether a large language model could do that job — and published the results in the Journal of the American College of Radiology, as reported by Radiology Business.

    The researchers retrospectively pulled breast ultrasound reports from 735 patients with pathologically confirmed breast masses, drawn from 60 hospitals across China. They compared an LLM — Qwen2.5-VL-7B — against hospital QC personnel who manually convert free-text reports into standardized, BI-RADS-based structured reports. Both accuracy and time-to-completion were measured.

    "Trained on extensive medical text corpora, LLMs have demonstrated proficiency in understanding clinical terminology, contextual relationships, and logical reasoning," the authors wrote. "These strengths align closely with the requirements of report [quality control], enabling LLMs to automatically identify errors, inconsistencies, and omissions in US reports in accordance with established reporting standards such as BI-RADS."

    Where the LLM beat manual review — and where it held up

    On two of the lesion characteristics reported, the LLM outperformed the humans doing the same job: 80% accuracy versus 64% for lesion margins, and 74% versus 56% for echo patterns. That gap didn't shrink on harder cases — the model's accuracy held up on reports describing multiple lesions, and it actually improved as BI-RADS categories rose from 3 to 5, the range covering more clinically suspicious findings.

    "Importantly, even in the more complex context of multi-lesion reports, the LLM demonstrated accuracy comparable to manual [quality control], highlighting its potential for processing complex clinical narratives," the authors noted. "This advantage likely stems from the LLM's extensive pretraining on medical texts, strong contextual understanding, and resistance to fatigue, ensuring stable and consistent QC performance."

    The time math is the real headline

    Accuracy gains are notable, but the workflow number is what should get a department's attention: the LLM completed quality control on 50 reports in an average of 13 minutes, versus 213 minutes for manual reviewers doing the same conversion and check. That's roughly a 16-fold difference, in a task departments already have to do — it doesn't require choosing to add a new review step, only automating one that already exists.

    MeasureQwen2.5-VL-7B (LLM)Manual QC reviewers
    Margin accuracy80%64%
    Echo pattern accuracy74%56%
    Time for 50 reports13 minutes213 minutes
    Performance on multi-lesion reportsHeld steadyReference

    An audit layer is a different job than drafting

    It's worth being precise about what this study evaluated. This is quality control applied to already-completed reports — checking finished text for errors, inconsistencies, and standards compliance after the fact. That's a different job from an AI producing the first draft of a report before a radiologist ever reviews it.

    The two aren't competing approaches; they're complementary safety layers that catch different things at different points. A retrospective LLM audit can flag drift in reporting standards across a department, a shift, or a hospital network. A draft-then-review workflow puts a check in front of the report before it's finalized at all, so fewer errors reach the audit stage in the first place. Departments evaluating AI tools for quality don't have to pick one over the other — the research suggests both are increasingly viable.

    Where xAID fits

    xAID's model puts a review step earlier in the process rather than relying solely on a downstream audit: the foundation model produces a structured, comprehensive report draft, xAID's in-house radiologist reviews every preliminary, and the report reaches the client ready-to-sign. This QC research reinforces the same underlying premise that draft-then-sign reporting is built on — that current-generation language models are already accurate and consistent enough to be trusted as a genuine check on report quality, not just a novelty. Whether that check happens as a pre-signature review or a post-hoc audit, the direction of travel is the same: AI narrows where human attention needs to go, it doesn't remove the human from the loop.

    The authors themselves are careful not to overclaim. They note open questions before this scales beyond a research setting: data privacy safeguards, the need to keep models updated as reporting standards evolve, and real-world testing of user acceptance, infrastructure requirements, and long-term stability. It's promising evidence, not a finished product.

    Frequently asked questions

    What did the new study find about LLMs and radiology quality control?

    A retrospective study published in the Journal of the American College of Radiology tested an LLM (Qwen2.5-VL-7B) against hospital quality-control staff on breast ultrasound reports from 735 patients with pathologically confirmed breast masses across 60 hospitals in China. The LLM matched or exceeded manual reviewers on accuracy for key BI-RADS reporting elements and did it far faster.

    How much faster was the LLM than manual radiology quality control review?

    The LLM completed quality control on 50 reports in an average of 13 minutes, compared with 213 minutes for manual reviewers converting the same free-text reports into standardized BI-RADS structured reports — roughly 16 times faster.

    Was the LLM more accurate than human reviewers at radiology quality control?

    On the two lesion characteristics reported, yes: the LLM scored 80% accuracy on margins versus 64% for manual reviewers, and 74% versus 56% on echo patterns. Its accuracy held up on more complex, multi-lesion reports and improved as BI-RADS categories rose from 3 to 5 — the more clinically suspicious findings.

    Does this mean AI can replace radiologists in quality assurance?

    No — the study's authors frame it as a support tool, not a replacement, and flag open questions around data privacy, model updating, and real-world stability before wider deployment. It's evidence that LLMs can serve as a second-layer check on report quality, complementing — not replacing — radiologist review of the report itself.

    Source: study published in the Journal of the American College of Radiology (2026), available via ScienceDirect, as reported by Radiology Business. Figures are rounded as reported.

    Review, built in from the first draft.

    xAID's foundation model drafts a comprehensive report, xAID's in-house radiologist reviews every preliminary, and your radiologist gets it ready-to-sign. Try it on 5 free studies.