构建全球多模态医学影像推理数据集,助力AI更懂医生怎么看片。
CheXthought: A global multimodal dataset of clinical chain-of-thought reasoning and visual attention for chest X-ray interpretation

- 采集501位放射科医生对5万张胸片的思考过程与注视轨迹。
- 用该数据训练的模型在定位准确性和错误减少上显著优于现有方法。
- 适合研究医疗AI可解释性、人机协作与诊断一致性的人群。
胸部X光解读是临床中最常见的诊断任务之一,也是AI研究的重点方向。然而,现有视觉-语言模型主要基于图像与报告配对数据训练,缺乏对临床推理认知过程和视觉注意力机制的建模。本文提出CheXthought,一个全球性的多模态数据集,包含来自71个国家501名放射科医生对50,312张多读胸片的103,592条链式思维推理记录和6,609,082个同步视觉注意力标注。分析揭示了专家在不同视觉搜索策略、临床信息整合及不确定性表达方面的认知模式。实验证明:第一,基于CheXthought的推理显著优于当前最优视觉-语言模型,在事实准确性和空间定位上表现更好;第二,将视觉注意力作为推理时提示可有效发现遗漏病灶并大幅降低幻觉;第三,使用该数据训练的模型在病灶分类、视觉忠实度、时间推理和不确定性表达方面均有显著提升;第四,利用多阅片者标注,可直接从图像预测人类间及人机分歧,实现病例难度、不确定性和模型可靠性透明化。这为推进多模态临床推理和开发更透明、可解释的视觉-语言模型提供了重要资源。
原文摘要 · Abstract (English)
Chest X-ray interpretation is one of the most frequently performed diagnostic tasks in medicine and a primary target for AI development, yet current vision-language models are primarily trained on datasets of paired images and reports, not the cognitive processes and visual attention that underlie clinical reasoning. Here, we present CheXthought, a global, multimodal resource containing 103,592 chain-of-thought reasoning traces and 6,609,082 synchronized visual attention annotations across 50,312 multi-read chest X-rays from 501 radiologists in 71 countries. Our analysis reveals clinical reasoning patterns in how experts deploy distinct visual search strategies, integrate clinical context, and communicate uncertainty. We demonstrate the clinical utility of CheXthought across four dimensions. First, CheXthought reasoning significantly outperforms state-of-the-art vision-language model chain-of-thought in factual accuracy and spatial grounding. Second, visual attention data used as an inference-time hint recovers missed findings and significantly reduces hallucinations. Third, vision-language models trained on CheXthought data achieve significantly stronger pathology classification, visual faithfulness, temporal reasoning and uncertainty communication. Fourth, leveraging CheXthought's multi-reader annotations, we predict both human-human and human-AI disagreement directly from an image, enabling transparent communication of case difficulty, uncertainty and model reliability. These findings establish CheXthought as a resource for advancing multimodal clinical reasoning and the development of more transparent, interpretable vision-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。