arXiv:2602.21657cs.CVcs.AI2026-02被引 1

让AI学习医生看片时的视线轨迹,提升胸部X光诊断的可解释性与协作性。

Following the Diagnostic Trace: Visual Cognition-guided Cooperative Network for Chest X-Ray Diagnosis

  • 通过眼动追踪捕捉医生看片时的注意力路径,引导AI学习人类视觉认知模式。
  • 在三个数据集上准确率达88.4%至92.4%,注意力分布与医生眼动高度一致。
  • 适合医学AI研发者、放射科医生及人机协同诊断系统设计者参考。

计算机辅助诊断(CAD)虽推动了胸部X光自动化诊断的发展,但仍与临床流程脱节,缺乏可靠决策支持和可解释性。人机协作旨在通过整合可控放射科医生的行为来提升诊断模型可靠性,但缺乏嵌入诊断流程的交互工具,以及医生决策模式与模型表征间的语义鸿沟,制约了临床应用。为此,本文提出视觉认知引导的协作网络(VCC-Net),实现协同诊断范式。VCC-Net以视觉认知(VC)为核心,利用眼动追踪或鼠标操作等临床兼容接口,捕获医生诊断过程中的视觉搜索轨迹与注意力模式。通过将视觉认知作为空间认知引导,学习分层视觉搜索策略,定位诊断关键区域。随后,认知图协同编辑模块融合医生视觉认知与模型推理,构建疾病感知图谱,捕捉解剖区域间依赖关系,对齐模型表征与认知驱动特征,缓解医生偏见,促进互补且透明的决策。在公开数据集SIIM-ACR、EGD-CXR及自建TB-Mouse数据集上,分类准确率分别达88.40%、85.05%和92.41%。VCC-Net生成的注意力图与医生眼动分布高度吻合,体现医生与模型推理的相互增强。代码已开源:https://github.com/IPMI-NWU/VCC-Net。

原文摘要 · Abstract (English)

Computer-aided diagnosis (CAD) has significantly advanced automated chest X-ray diagnosis but remains isolated from clinical workflows and lacks reliable decision support and interpretability. Human-AI collaboration seeks to enhance the reliability of diagnostic models by integrating the behaviors of controllable radiologists. However, the absence of interactive tools seamlessly embedded within diagnostic routines impedes collaboration, while the semantic gap between radiologists' decision-making patterns and model representations further limits clinical adoption. To overcome these limitations, we propose a visual cognition-guided collaborative network (VCC-Net) to achieve the cooperative diagnostic paradigm. VCC-Net centers on visual cognition (VC) and employs clinically compatible interfaces, such as eye-tracking or the mouse, to capture radiologists' visual search traces and attention patterns during diagnosis. VCC-Net employs VC as a spatial cognition guide, learning hierarchical visual search strategies to localize diagnostically key regions. A cognition-graph co-editing module subsequently integrates radiologist VC with model inference to construct a disease-aware graph. The module captures dependencies among anatomical regions and aligns model representations with VC-driven features, mitigating radiologist bias and facilitating complementary, transparent decision-making. Experiments on the public datasets SIIM-ACR, EGD-CXR, and self-constructed TB-Mouse dataset achieved classification accuracies of 88.40%, 85.05%, and 92.41%, respectively. The attention maps produced by VCC-Net exhibit strong concordance with radiologists' gaze distributions, demonstrating a mutual reinforcement of radiologist and model inference. The code is available at https://github.com/IPMI-NWU/VCC-Net.

医学影像人机协作可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。