arXiv:2509.06409cs.AI2025-09被引 1

用报告引导思维链,让AI学会像医生一样逐步诊断。

Teaching AI Stepwise Diagnostic Reasoning with Report-Guided Chain-of-Thought Learning

  • 用报告引导思维链训练,模仿医生分步推理过程。
  • 在MIMIC-CXR上疾病分类AUC提升至0.76,准确率大幅提高。
  • 适合想开发可解释医疗AI的研究者与临床工程师。

本研究提出DiagCoT多阶段框架,通过监督微调通用视觉语言模型(VLMs),仅利用自由文本报告即可模拟放射科医生的分步诊断推理。DiagCoT结合对比图像-报告调优实现领域对齐,链式思维监督捕捉推理逻辑,并采用带临床奖励信号的强化调优提升事实准确性与语言流畅性。在MIMIC-CXR基准测试中,DiagCoT将零样本疾病分类AUC从0.52提升至0.76(绝对提升0.24),病理定位mIoU从0.08升至0.31(绝对提升0.23),报告生成BLEU从0.11增至0.33(绝对提升0.22)。其在长尾疾病和外部数据集上均优于当前最优模型如LLaVA-Med和CXR-LLAVA。通过将非结构化临床叙述转化为结构化监督信号,DiagCoT为构建可解释且具备诊断能力的医疗AI系统提供了可扩展方案。

原文摘要 · Abstract (English)

This study presents DiagCoT, a multi-stage framework that applies supervised fine-tuning to general-purpose vision-language models (VLMs) to emulate radiologists' stepwise diagnostic reasoning using only free-text reports. DiagCoT combines contrastive image-report tuning for domain alignment, chain-of-thought supervision to capture inferential logic, and reinforcement tuning with clinical reward signals to enhance factual accuracy and fluency. On the MIMIC-CXR benchmark, DiagCoT improved zero-shot disease classification AUC from 0.52 to 0.76 (absolute gain of 0.24), pathology grounding mIoU from 0.08 to 0.31 (absolute gain of 0.23), and report generation BLEU from 0.11 to 0.33 (absolute gain of 0.22). It outperformed state-of-the-art models including LLaVA-Med and CXR-LLAVA on long-tailed diseases and external datasets. By converting unstructured clinical narratives into structured supervision, DiagCoT offers a scalable approach for developing interpretable and diagnostically competent AI systems for radiology.

医疗AI思维链视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。