What the study actually measured
The analysis drew on the high-resolution CT substudy of the Phase III INBUILD trial, a double-blind, randomized, placebo-controlled study of the antifibrotic drug nintedanib in patients with progressive pulmonary fibrosis (PPF) — interstitial lung diseases, other than idiopathic pulmonary fibrosis, that keep worsening despite treatment. Researchers applied two independent quantitative CT methods to baseline, 24-week, and 52-week scans from 474 patients, publishing the results in the American Journal of Respiratory and Critical Care Medicine.
Both methods scored the same scans for extent of fibrotic disease. One measure of quantitative fibrosis extent showed roughly a 7% relative difference favoring the treated group at 24 weeks (95% CI -11 to -2; p=0.005), persisting through 52 weeks. A second, independently developed total-disease-extent score showed a similar pattern — roughly an 8% relative difference at 24 weeks (95% CI -12 to -4; p<0.001) and about 7% at 52 weeks (p<0.05). Related reticulovascular scoring also showed a significant treatment effect at 24 weeks, per the same substudy.
The two independently built scoring methods — one validated against a UCLA research algorithm — moved together, which is itself notable: it means the signal isn't an artifact of one vendor's math. It's a property of the scans.
The finding underneath the finding
It's tempting to read this as a story about a clever new algorithm. It's really a story about data that already existed. Every patient in this substudy had the same chest CT scans a routine PPF monitoring protocol would order anyway — at baseline, 24 weeks, and 52 weeks. Nothing about the scanning changed. What changed was whether anyone systematically measured, scored, and compared what the images showed over time.
That's the detail easy to miss in a vendor press cycle: the researchers didn't discover a new biomarker hiding somewhere exotic. They extracted a quantitative, longitudinal signal — disease extent and reticulovascular pattern, tracked scan-over-scan — from the same CT data radiologists have been reading for this disease for years. The substudy also found that, in the placebo arm, patients with higher baseline scores had a faster rate of forced vital capacity (FVC) decline over 52 weeks, with restricted mean survival time differences of roughly 40 to 65 days between patients above and below the median score. That's prognostic information sitting in the imaging, waiting to be captured consistently.
None of that requires a new scanner, a new contrast protocol, or a new patient visit. It requires turning an image into a number — the same number, scored the same way, every time — instead of into a paragraph of prose that varies by who dictated it and how they were feeling that day.
Why narrative reports throw most of that signal away
A standard chest CT report for interstitial lung disease is dictated narrative text: "extent of fibrosis appears stable compared to prior," "mild interval progression of reticulation," "ground-glass opacities grossly unchanged." Useful shorthand for a single read — but none of it is a number, and none of it is reliably comparable to the same radiologist's own wording six months later, let alone a different radiologist's.
That's the bottleneck this research actually exposes. The imaging already generates the longitudinal, quantifiable data that trial-grade quantitative CT scoring extracted. A dictated narrative report just isn't built to carry it forward — there's no structured field for "percent disease extent," no persistent score a clinic can trend visit over visit, no reproducible number a payer or a tumor board (or in this case, a fibrosis clinic) can act on without re-reading every prior scan by eye.
| What's needed to track PPF treatment response | Narrative dictated report | Structured, quantitative report |
|---|---|---|
| Disease extent as a number | Rarely — described in words ("mild," "extensive") | Scored and recorded each read |
| Comparable across visits | Depends on consistent wording by the same reader | Same metric, same scale, every scan |
| Trendable over 24–52 weeks | Requires manually re-reading prior reports | Plotted automatically as a series |
| Usable as a trial-grade endpoint | No | Yes — as this substudy demonstrates |
Where this fits with how AI CT reporting actually works
This is the clinical-trial version of a gap that shows up in everyday reads across every organ system, not just the lung: imaging already contains quantifiable, comparable detail that narrative dictation compresses into prose and discards. AI CT reporting built on a foundation model is designed to close exactly that gap at the point of care — producing a structured, comprehensive report draft with measurements captured consistently rather than described loosely, so findings stay comparable from one scan to the next. xAID's in-house radiologist reviews every preliminary before it reaches the client's reading radiologist, who signs the final — the structure doesn't replace clinical judgment, it gives that judgment a consistent number to work from.
Frequently asked questions
Can a pulmonary fibrosis CT scan measure whether treatment is working?
A 2026 analysis of 474 patients from the Phase III INBUILD trial's high-resolution CT substudy, published in the American Journal of Respiratory and Critical Care Medicine, found that quantitative CT scoring detected a treatment effect from the antifibrotic nintedanib as early as 24 weeks and sustained through 52 weeks. One quantitative extent-of-fibrosis score showed roughly a 7% relative difference favoring treatment at 24 weeks (95% CI -11 to -2; p=0.005), and a second measure of total disease extent showed roughly an 8% relative difference (95% CI -12 to -4; p<0.001).
What is progressive pulmonary fibrosis (PPF)?
Progressive pulmonary fibrosis describes interstitial lung diseases, other than idiopathic pulmonary fibrosis, that worsen over time despite treatment — with increasing scarring, declining lung function, and a prognosis that can approach that of IPF. The INBUILD trial established nintedanib as an approved antifibrotic treatment for this group.
Why don't standard CT reports capture quantitative treatment response?
Routine chest CT reports for interstitial lung disease are typically dictated narrative text — terms like 'stable' or 'mild progression' rather than a reproducible numeric score. The INBUILD substudy shows the underlying CT data already contains a quantifiable, trackable signal; the bottleneck is that unstructured reporting doesn't systematically extract and record it in a comparable form across scans.
Does quantitative CT predict lung function decline in pulmonary fibrosis?
Yes. In the same substudy, within the placebo arm, patients with higher baseline quantitative CT scores had a faster rate of decline in forced vital capacity (FVC) over 52 weeks, with restricted mean survival time differences of roughly 40 to 65 days between patients above versus below the median baseline score.
Source: Devaraj A, et al. "Quantitative Computed Tomography in Progressive Pulmonary Fibrosis: Data from a Sub-Study of the Double Blind, Randomized, Placebo-controlled INBUILD Trial," American Journal of Respiratory and Critical Care Medicine (2026), doi.org/10.1093/ajrccm/aamag526. Figures are rounded as reported.