← BlogPatient Safety & QAJuly 25, 20267 min read

    Why the words radiologists use
    can delay a patient's care

    A new study of nearly 1,000 CT scans found that reclassifying reports by wording alone — no images re-reviewed — cut the indeterminate rate substantially. Radiology report language isn't a style choice. It's a variable that changes what happens to the patient next.

    15.5%
    Original CT reports classified indeterminate
    suspected appendicitis, n=983
    12%
    Indeterminate rate after wording alone was re-scored
    no images re-reviewed
    61.3%
    Indeterminate-report patients who had surgery
    vs. 3.1% with negative reports
    983
    CT exams analyzed
    Aug 2021 – Jul 2025

    What the study looked at

    Researchers at Siriraj Hospital in Bangkok, led by Piyachai Siriphiphatcharoen with Rathachai Kaewlai as corresponding author, published a retrospective analysis in Insights into Imaging examining how radiology report wording affects the diagnosis of acute appendicitis on CT — one of the most common reasons a CT is ordered in an emergency department. The team reviewed 983 CT exams for suspected appendicitis performed at a 2,200-bed academic medical center between August 2021 and July 2025, issued by 39 board-certified radiologists.

    They compared four ways of expressing diagnostic certainty on the same set of scans: the original radiologist report as dictated; an adjudicated version, where two radiologists re-scored the same report's wording against a simplified three-tier framework (negative, indeterminate, positive) without looking at the images again; and two rule-based lexicon schemes that graded appendicitis likelihood purely from CT measurements — appendiceal diameter, wall thickening, wall hyperenhancement, and periappendiceal fat stranding — using either a 6 mm or 7 mm diameter cutoff. Surgical, pathology, and clinical outcomes served as the reference standard. As the authors put it: "When CT findings are inconclusive, clinical decision-making becomes more complex, and management pathways may vary considerably across providers."

    The finding that matters: language, not the image, drove much of the uncertainty

    Acute appendicitis was ultimately confirmed in 40% of the 983 patients. The original reports were classified as indeterminate in 15.5% of cases. When the two adjudicating radiologists re-scored only the wording of those same reports against the three-tier framework — without re-reading a single image — the indeterminate rate fell to 12%, entirely by reclassifying ambiguous phrasing as negative.

    The authors were explicit about what that means: "The fact that reclassification was possible through wording reassessment alone suggests that a meaningful proportion of indeterminacy in our cohort reflected imprecise language rather than unavoidable diagnostic ambiguity." Put plainly, a chunk of the uncertainty patients and referrers were dealing with wasn't coming from the scan — it was coming from how the finding was written down.

    Indeterminate language has a real downstream cost

    This is where the study moves from a linguistics curiosity to a patient-safety issue. Among patients whose report was indeterminate, 61.3% went on to have surgery for presumed appendicitis — compared with just 3.1% of patients whose report was negative. Indeterminate reports were also far more likely to carry a generic "clinical correlation" caveat: 32.8% of indeterminate reports included that phrase, versus 6.3% of negative reports and only 0.9% of positive ones.

    Hedging language doesn't just create an "unanswered question," as the researchers describe it — it pushes clinical decision-making onto a referring physician who has less information than the radiologist looking at the scan, and it can tilt that decision toward intervention. The study's own reference-standard data show that even among indeterminate cases, appendicitis was confirmed only 35.5% of the time — far below the 61.3% surgical rate. Some of that gap is unavoidable clinical caution. Some of it, per the authors' own reclassification, was avoidable imprecision.

    Standardized categories help — but a rulebook alone isn't the answer

    It would be tempting to conclude that a rigid, checklist-style lexicon — the kind BI-RADS or LI-RADS uses in other domains — is the fix. The data don't support that. The two rule-based lexicon schemes in this study, which scored appendicitis likelihood purely from CT measurements, produced higher indeterminate rates than either the original or the adjudicated reports: 27.9% for the 7 mm cutoff and 46.1% for the 6 mm cutoff. Agreement with the final diagnosis (Cohen's kappa) also came in lower for both lexicons than for the radiologist-driven approaches.

    The best outcome in the study came from a third option: a radiologist's full clinical judgment, expressed through a small, defined set of certainty categories rather than free-text hedging. That's consistent with a separate 2025 paper from the American College of Radiology's Commission on Quality and Safety, which proposed standardized frameworks for communicating diagnostic certainty precisely because, as it notes, there is no universal terminology system for expressing uncertainty in a radiology report today.

    Reporting approachIndeterminate rateAgreement with final diagnosis (kappa)What drives the result
    Original report (free-text)15.5%0.759–0.838Real judgment, expressed in ad hoc hedging language
    Rule-based lexicon (7mm cutoff)27.9%0.562–0.796Mechanical CT-measurement checklist, no contextual weighting
    Rule-based lexicon (6mm cutoff)46.1%0.292–0.784Narrower cutoff creates more borderline calls
    Standardized 3-tier wording (adjudicated, same images)12%0.819–0.838Same clinical judgment, precise and consistent language

    Why this matters beyond one diagnosis

    Appendicitis is a useful test case because it has a clean reference standard — surgery, pathology, and clinical follow-up all confirm or refute the call. Most findings radiologists report don't have that luxury, which is exactly why the wording matters more, not less: a referring physician acting on a report for a lung nodule, a liver lesion, or a pulmonary embolus rule-out doesn't get a same-day answer the way an appendix does. Three implications follow for any imaging organization:

    Indeterminate is a real category, not a failure to decide

    The study treats "indeterminate" as a distinct, legitimate diagnostic outcome rather than a radiologist punting on a call. The goal isn’t zero indeterminate reports — it’s making sure indeterminate is used only when the finding is genuinely ambiguous, not when the wording is.

    A small, defined vocabulary beats an open-ended one

    Reducing the indeterminate rate required only a shared three-tier vocabulary, applied consistently by more than one reader. That is a language and process fix, not a new imaging protocol or a longer read time.

    Checklists support judgment; they don’t replace it

    The lexicon-based schemes in this study underperformed radiologist judgment precisely because they couldn’t weigh findings in context. Structured categories work best layered onto a full clinical read, not substituted for one.

    Where this fits with AI-drafted reporting

    This is the case for structured, foundation-model-drafted reports built around a small, defined set of diagnostic-certainty categories rather than free-text hedging. A report draft can apply the same certainty threshold to every study on every shift, so "correlate clinically" doesn't become a stand-in for a judgment the report never actually renders — and the language stays consistent whether the case comes in at 2 p.m. or 2 a.m. That structured draft still goes through an in-house radiologist review before it reaches the client's reading radiologist ready-to-sign; precise, standardized language is a complement to that review, not a replacement for it.

    Frequently asked questions

    What did the new study find about radiology report language?

    A 2026 study in Insights into Imaging, led by researchers at Siriraj Hospital in Bangkok, examined 983 CT scans performed for suspected acute appendicitis between 2021 and 2025. It found that 15.5% of the original radiology reports were classified as indeterminate, and that re-assessing the wording alone, without looking at the images again, resolved enough of those cases to cut the indeterminate rate to 12%.

    How much did standardizing report wording reduce indeterminate results?

    Re-scoring report wording against a simplified three-tier certainty framework (negative, indeterminate, positive) reduced the indeterminate rate from 15.5% to 12%, a reduction achieved purely by tightening language, since no CT images were re-reviewed. The two radiologists doing the re-scoring agreed with each other in the large majority of cases, with a Cohen's kappa of 0.711.

    Does vague radiology report language actually affect patient care?

    Yes. In the same study, 61.3% of patients whose CT report was indeterminate went on to have surgery for presumed appendicitis, compared with just 3.1% of patients whose report was negative. Indeterminate reports were also far more likely to include a generic 'clinical correlation' caveat (32.8%) than negative (6.3%) or positive (0.9%) reports, a sign that imprecise language, not only genuine diagnostic difficulty, was steering downstream decisions.

    Is a BI-RADS-style checklist the fix for ambiguous radiology reports?

    Not on its own, according to the study. Two rule-based lexicon schemes that graded appendicitis likelihood purely from CT measurements produced higher indeterminate rates (27.9% and 46.1%) than either the original or the standardized radiologist reports. The best results came from combining a radiologist's full clinical judgment with a simplified, standardized certainty framework, not from replacing judgment with a checklist.

    Source: Siriphiphatcharoen P, Kaewlai R, et al. "Indeterminate CT reporting of adult acute appendicitis: radiologist versus lexicon-based diagnostic certainty." Insights into Imaging (2026), DOI: 10.1186/s13244-026-02352-y, as reported by Radiology Business. Framework context: Shinagare AB, et al. "Potential Frameworks for Communicating Diagnostic Certainty in Radiology Reports: From the ACR Commission on Quality and Safety." J Am Coll Radiol 22:1390–1398 (2025), DOI: 10.1016/j.jacr.2025.07.027. Figures are rounded as reported.

    One structured, consistent report draft. Every study.

    xAID's foundation model applies the same certainty language on every case, and an in-house radiologist reviews every preliminary before it's ready for your radiologist to sign. Try it on 5 free studies.