arXiv:2603.22939cs.CVcs.LG2026-03

用专家注视轨迹直接提升胸片分类准确率

FixationFormer: Direct Utilization of Expert Gaze Trajectories for Chest X-Ray Classification

  • 将注视轨迹当作序列输入Transformer,保留时空结构
  • 在3个公开数据集上达到当前最优分类效果
  • 适合研究医学影像诊断推理与注意力机制结合的学者

专家眼动提供了放射科领域丰富的被动知识,是融合诊断推理到辅助分析中的有力线索。然而,传统基于CNN的医疗图像分析系统难以直接整合注视数据:注视记录具有时序密集、空间稀疏、噪声大且专家间差异显著的特点。因此,现有方法多采用热力图等简化表示。相比之下,注视数据与Transformer架构天然契合,二者均为序列结构并依赖注意力机制聚焦关键区域。本文提出FixationFormer,一种基于Transformer的架构,将专家注视轨迹表示为令牌序列,从而保留其时空结构。通过联合建模图像特征与注视序列,该方法缓解了注视数据的稀疏性与变异性,实现诊断线索的细粒度直接整合。我们在三个公开胸片数据集上验证该方法,结果表明其分类性能达到当前最优水平,证明了在Transformer框架中以序列形式表示注视轨迹的价值。

原文摘要 · Abstract (English)

Expert eye movements provide a rich, passive source of domain knowledge in radiology, offering a powerful cue for integrating diagnostic reasoning into computer-aided analysis. However, direct integration into CNN-based systems, which historically have dominated the medical image analysis domain, is challenging: gaze recordings are sequential, temporally dense yet spatially sparse, noisy, and variable across experts. As a consequence, most existing image-based models utilize reduced representations such as heatmaps. In contrast, gaze naturally aligns with transformer architectures, as both are sequential in nature and rely on attention to highlight relevant input regions. In this work, we introduce FixationFormer, a transformer-based architecture that represents expert gaze trajectories as sequences of tokens, thereby preserving their temporal and spatial structure. By modeling gaze sequences jointly with image features, our approach addresses sparsity and variability in gaze data while enabling a more direct and fine-grained integration of expert diagnostic cues through explicit cross-attention between the image and gaze token sequences. We evaluate our method on three publicly available benchmark chest X-ray datasets and demonstrate that it achieves state-of-the-art classification performance, highlighting the value of representing gaze as a sequence in transformer-based medical image analysis.

医学影像注意力机制眼动追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。