arXiv:2604.00493cs.CVcs.AI2026-04被引 4

CheXOne让肺部X光诊断模型会说理,提升可解释性与临床实用性。

A Reasoning-Enabled Vision-Language Foundation Model for Chest X-ray Interpretation

  • 用1470万条标注数据训练,结合指令微调与强化学习生成推理链条。
  • 零样本测试中超越现有模型,在17个任务上表现优异。
  • 生成的推理路径符合临床逻辑,助医生提效且可信赖。

胸部X光(CXRs)是全球最频繁的影像检查之一,但影像量激增加重放射科医生负担并增加误诊风险。尽管人工智能在CXR解读中展现潜力,多数系统仅输出最终结论,缺乏对视觉证据如何转化为诊断依据的明确解释。我们提出CheXOne,一种支持推理的视觉语言基础模型,可联合生成诊断预测与显式、临床相关的推理链条,连接视觉证据、影像发现与诊断结果。该模型基于从30个公开数据集整理的1470万条指令与推理样本,覆盖36种CXR解读任务,采用两阶段框架,结合指令微调与强化学习以提升推理质量。我们在零样本设置下评估了17个视觉问答、报告生成、视觉定位与推理评估场景,结果表明,CheXOne优于现有医学与通用领域基础模型,并在独立公共基准上表现强劲。临床读者研究显示,55%情况下,CheXOne生成的报告质量相当于或优于住院医师水平,有效回应临床指征,提升报告撰写与读片效率。进一步由放射科医生参与的分析表明,生成的推理链条具备高临床真实性,为最终预测提供因果支持,合理解释性能提升。这些结果表明,显式推理可增强模型性能、可解释性与临床应用价值。

原文摘要 · Abstract (English)

Chest X-rays (CXRs) are among the most frequently performed imaging examinations worldwide, yet rising imaging volumes increase radiologist workload and the risk of diagnostic errors. Although artificial intelligence (AI) systems have shown promise for CXR interpretation, most generate only final predictions, without making explicit how visual evidence is translated into radiographic findings and diagnostic predictions. We present CheXOne, a reasoning-enabled vision-language model for CXR interpretation. CheXOne jointly generates diagnostic predictions and explicit, clinically grounded reasoning traces that connect visual evidence, radiographic findings, and these predictions. The model is trained on 14.7 million instruction and reasoning samples curated from 30 public datasets spanning 36 CXR interpretation tasks, using a two-stage framework that combines instruction tuning with reinforcement learning to improve reasoning quality. We evaluate CheXOne in zero-shot settings across visual question answering, report generation, visual grounding and reasoning assessment, covering 17 evaluation settings. CheXOne outperforms existing medical and general-domain foundation models and achieves strong performance on independent public benchmarks. A clinical reader study demonstrates that CheXOne-drafted reports are comparable to or better than resident-written reports in 55% of cases, while effectively addressing clinical indications and enhancing both report writing and CXR interpretation efficiency. Further analyses involving radiologists reveal that the generated reasoning traces show high clinical factuality and provide causal support for the final predictions, offering a plausible explanation for the performance gains. These results suggest that explicit reasoning can improve model performance, interpretability and clinical utility in AI-assisted CXR interpretation.

医学AI视觉推理可解释性胸部X光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。