arXiv:2605.23992cs.CVcs.AI2026-05被引 2

用放射科医生注视轨迹建模图像阅读过程,提升医学影像诊断性能

A World Model of Radiologist Reading for Medical Image Representation Learning

论文配图:A World Model of Radiologist Reading for Medical Image Representation Learning
图 1 · 摘自论文原文
  • 将医生注视序列视为图像中的移动轨迹,自回归预测下一块区域的特征
  • 在三个数据集上达到最佳监督与零样本诊断准确率,零样本性能领先
  • 无需真实注视数据即可生成高质量图像表示,适合医学影像预训练

放射科医生的眼动追踪数据完整记录了专家在阅片时的搜索、比对和证据积累过程;然而现有方法仅部分利用这一信号,或作为静态空间先验,或作为与诊断解耦的辅助目标。我们提出 GazeWorld,一种医学影像世界模型,将图像视为世界,医生的注视序列视为其中的轨迹。GazeWorld 自回归地从已访问区域的特征中预测下一个注视块的潜在表示,同时通过空间补全分支覆盖未访问区域。推理时,GazeWorld 仅需图像即可生成完整的块特征序列,无需真实眼动数据。冻结的 GazeWorld 特征在 CheXpert、RSNA Pneumonia 与 SIIM-ACR Pneumothorax 的九种监督设置下均达到最先进水平,并在三者上实现最高零样本准确率。在 GazeSearch 基准测试中,基于相同冻结特征训练的通用解码器,在 ScanMatch 和 SED 上分别优于专门设计的 LogitGaze-Med 超过 16% 和 22%,尽管未显式训练预测注视。GazeWorld 表明,建模专家如何阅读,而不仅是他们得出的结论,为医学影像 AI 提供了有前景的预训练范式。

原文摘要 · Abstract (English)

Radiologist eye-tracking data provide a rich record of how experts search, compare, and accumulate evidence during image reading; yet, existing methods exploit this signal only partially, either as a static spatial prior or as an auxiliary prediction target decoupled from diagnosis. We propose GazeWorld, a medical imaging world model that treats the image as the world and the radiologist's fixation sequence as a trajectory through it. GazeWorld autoregressively predicts the latent representation of the next fixated patch from all previously visited ones, while a spatial-completion branch covers unvisited regions. At inference, GazeWorld generates a sequence of patch representations from the image alone without requiring real gaze data. Frozen GazeWorld features achieve state-of-the-art diagnostic accuracy across all nine supervised settings on CheXpert, RSNA Pneumonia, and SIIM-ACR Pneumothorax, as well as the highest zero-shot accuracy on all three benchmarks. On the GazeSearch benchmark, a generic decoder trained on the same frozen features outperforms the purpose-built LogitGaze-Med by over 16\% in ScanMatch and 22\% in SED, despite not being explicitly trained to predict gaze. GazeWorld demonstrates that modeling how experts read, not just what they conclude, offers a promising pretraining paradigm for medical imaging AI.

医学影像世界模型眼动追踪预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。