用思维链让AI像医生一样解释胸片诊断,结果更可信。
X-Ray-CoT: Interpretable Chest X-ray Diagnosis with Vision-Language Models via Chain-of-Thought Reasoning
- 模仿医生思考过程,结合视觉与语言模型推理
- 在CORDA数据集上达80.52%准确率和78.65%F1值
- 生成可解释的诊断报告,适合临床信任场景
胸部X光对肺部和心脏疾病诊断至关重要,但解读需丰富临床经验,且存在观察者间差异。深度学习模型虽精度高,但黑箱特性限制其在高风险医疗场景的应用。为此,我们提出X-Ray-CoT(胸片思维链),利用视觉-语言大模型实现智能胸片诊断与可解释报告生成。该框架通过提取多模态特征与视觉概念,再以结构化思维链提示策略驱动大语言模型进行推理,生成详细自然语言诊断报告。在CORDA数据集上评估显示,其疾病诊断的平衡准确率为80.52%,F1得分为78.65%,略优于现有黑箱模型。关键的是,其生成的报告质量高且具备可解释性,初步人评验证有效。消融实验确认了多模态融合与思维链推理的关键作用,证明其对鲁棒、透明医疗AI系统的重要意义。
原文摘要 · Abstract (English)
Chest X-ray imaging is crucial for diagnosing pulmonary and cardiac diseases, yet its interpretation demands extensive clinical experience and suffers from inter-observer variability. While deep learning models offer high diagnostic accuracy, their black-box nature hinders clinical adoption in high-stakes medical settings. To address this, we propose X-Ray-CoT (Chest X-Ray Chain-of-Thought), a novel framework leveraging Vision-Language Large Models (LVLMs) for intelligent chest X-ray diagnosis and interpretable report generation. X-Ray-CoT simulates human radiologists' "chain-of-thought" by first extracting multi-modal features and visual concepts, then employing an LLM-based component with a structured Chain-of-Thought prompting strategy to reason and produce detailed natural language diagnostic reports. Evaluated on the CORDA dataset, X-Ray-CoT achieves competitive quantitative performance, with a Balanced Accuracy of 80.52% and F1 score of 78.65% for disease diagnosis, slightly surpassing existing black-box models. Crucially, it uniquely generates high-quality, explainable reports, as validated by preliminary human evaluations. Our ablation studies confirm the integral role of each proposed component, highlighting the necessity of multi-modal fusion and CoT reasoning for robust and transparent medical AI. This work represents a significant step towards trustworthy and clinically actionable AI systems in medical imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。