arXiv:2603.06697cs.CVcs.AI2026-03被引 1

用眼动轨迹监督医学视觉模型,让其像医生一样逐步看图推理。

Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs

  • 引入专属注视令牌,按时间顺序预测眼动选中的图像区域。
  • 在MIMIC-EYE等数据集上显著提升诊断准确率,跨域泛化能力更强。
  • 适合研究医疗视觉推理、可解释AI或眼动分析的学者参考。

视觉-语言模型(VLM)将图像处理为视觉标记,但其推理过程常依赖文本,对需视觉引导的放射学任务不理想。放射科医生则通过序列化视觉搜索进行诊断;眼动追踪能记录这一过程,生成时间有序的注视轨迹,揭示证据获取的动态过程。本文利用眼动作为监督信号,引入少量专用注视标记,训练模型按时间顺序预测被注视的图像区域,促使模型模仿人类证据获取与整合方式。在MIMIC-EYE及多个外部零样本基准上的实验表明,该方法持续优于基线,实现领域内最佳性能,并提升跨域鲁棒性。结果表明,时序注视轨迹是学习视觉化医学推理的有效监督信号。

原文摘要 · Abstract (English)

Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential visual search; eye-tracking captures this process as time-ordered gaze trajectories that reveal how evidence is acquired over time. We use eye-gaze as supervision to guide VLM reasoning by introducing a small set of dedicated gaze tokens. These tokens are trained to predict gaze-selected image patch indices in temporal order, encouraging the model to follow human-like evidence acquisition and integration. Experiments on MIMIC-EYE and multiple external zero-shot benchmarks show consistent gains over baselines, achieving state-of-the-art in-domain performance and improved out-of-domain robustness. These results highlight temporally ordered gaze as an effective supervision signal for learning visually grounded medical reasoning.

视觉推理眼动追踪医学AI可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。