arXiv:2508.13068cs.CVcs.LG2025-08被引 3

用眼动数据训练胸部X光诊断模型,让报告更准且可解释。

Eyes on the Image: Gaze Supervised Multimodal Learning for Chest X-ray Diagnosis and Report Generation

  • 融合图像块、框选区域和医生注视点,用眼动信号监督模型注意力
  • 诊断准确率提升4.4%(AUC),报告与医生关注区域相关性达0.306
  • 适合需要可解释性医疗AI的临床研究者和开发者

医学视觉-语言模型仍难以模拟放射科医生的注意力分布,也难生成带有明确空间定位的报告。本文基于MIMIC-Eye数据集,提出两阶段多模态框架。第一阶段引入注视点-标记分类器,融合图像块、边界框掩码、文本嵌入与放射科医生眼动轨迹,采用课程调度的可信度校准复合损失函数进行监督,显著提升准确率与空间对齐性:加入注视点监督后,AUC提高4.4%,F1提升13.3%,皮尔逊相关系数达0.306,证实关注区域具有临床意义。第二阶段将分类结果转化为区域特定的诊断语句:通过置信度加权提取关键词,经专家词典映射至17个胸腔区域,并由提示式大模型扩展生成,相较关键词基线,临床术语的BERTScore与ROUGE分数均显著提升。所有模块可开关消融,全流程可复现,为可解释、注视感知的胸部X光分析提供新基准。眼动信号的引入明显提升了诊断准确性和报告透明度。

原文摘要 · Abstract (English)

Medical vision-language models still struggle to match radiologists' attention and to verbalize findings with explicit spatial grounding. We address this gap with a two-stage multimodal framework for chest X-ray interpretation built on the MIMIC-Eye dataset. In the first stage introduces a gaze-token classifier that fuses image patches, bounding-box masks, transcription embeddings, and radiologist fixations. A curriculum-scheduled, trust-calibrated composite loss supervises the gaze token, boosting both accuracy and spatial alignment. Adding fixation supervision raises AUC 4.4% and F1 13.3%, and Pearson correlation rises to 0.306, confirming clinically relevant focus. In stage 2, classifier predictions are translated into region-specific diagnostic sentences. Confidence-weighted keywords are extracted, mapped to 17 thoracic regions through an expert dictionary, and expanded with a prompted large language model, boosting clinical-term BERTScore and ROUGE scores over keyword baselines. All components are toggle-able for ablation, and the full pipeline is reproducible, offering a new benchmark for interpretable, gaze-aware chest-X-ray analysis. Integrating eye-tracking signals demonstrably enhances both diagnostic accuracy and the transparency of generated reports.

医学影像多模态可解释性眼动追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。