用放射科医生的视线数据训练模型,让AI看X光片更像专家。
Seeing Through Experts Eyes A Foundational Vision Language Model Trained on Radiologists Gaze and Reasoning

- 用放射科医生眼动数据作为行为先验,学习专家看片顺序。
- 在30万张胸片上表现更准,与专家判断一致率达87%。
- 输出带可验证的视线轨迹和定位区域,适合临床协作使用。
大规模视觉语言模型在自动解读胸部X光片方面展现出潜力,但其临床实用性受限于模型输出与放射科医生推理过程之间的差距。现有系统多聚焦语义信息,未模拟专家如何观察医学图像,常遗漏关键发现或偏离标准诊断流程。放射科医生遵循结构化检查协议(如ABCDEF方法),确保所有临床相关区域被系统性评估,降低漏诊风险并支持可靠诊断推理。本文提出GazeX,一种利用放射科医生眼动数据作为行为先验的视觉语言模型,通过引入注视轨迹与凝视模式进行预训练,使模型学习专家注意力的空间与时间结构,并按临床有意义顺序整合观察结果。基于包含5名放射科医生超过3万帧注视关键帧的精选数据集,我们证明GazeX在报告生成、疾病定位和视觉问答任务中均实现更高准确率、更强可解释性及与专家的一致性。该模型基于231,835份影像学研究、780,014个问答对及1,162对带边界框的图像-句子对进行训练。与自主报告系统不同,GazeX生成可验证的证据产物,包括检查轨迹与关联的局部区域,支持高效人工核查与安全的人机协作。通过学习专家视角,为构建更可信、可解释且诊断稳健的医学AI系统提供可行路径。
原文摘要 · Abstract (English)
Large scale vision language models have shown promise in automating chest Xray interpretation, yet their clinical utility remains limited by a gap between model outputs and radiologist reasoning. Most systems optimize for semantic information without emulating how experts visually examine medical images, often overlooking critical findings or diverging from established diagnostic workflows. Radiologists follow structured protocols (e.g., the ABCDEF approach) that ensure all clinically relevant regions are systematically examined, reducing missed findings and supporting reliable diagnostic reasoning. We introduce GazeX, a vision language model that leverages radiologists' eye tracking data as a behavioral prior to model expert diagnostic reasoning. By incorporating gaze trajectories and fixation patterns into pretraining, GazeX learns to follow the spatial and temporal structure of radiologist attention and integrates observations in a clinically meaningful sequence. Using a curated dataset of over 30,000 gaze key frames from five radiologists, we demonstrate that GazeX produces more accurate, interpretable, and expert consistent outputs across radiology report generation, disease grounding, and visual question answering, utilizing 231,835 radiographic studies, 780,014 question answer pairs, and 1,162 image sentence pairs with bounding boxes. Unlike autonomous reporting systems, GazeX produces verifiable evidence artifacts, including inspection trajectories and finding linked localized regions, enabling efficient human verification and safe human AI collaboration. Learning through expert eyes provides a practical route toward more trustworthy, explainable, and diagnostically robust AI systems for radiology and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。