arXiv:2607.26542cs.CV2026-07

用眼球轨迹提升医疗图像弱监督分割精度

From Spatial Semantics to Temporal Context: Leveraging Gaze Trajectory for Weakly Supervised Medical Image Segmentation

论文配图:From Spatial Semantics to Temporal Context: Leveraging Gaze Trajectory for Weakly Supervised Medical Image Segmentation
图 1 · 摘自论文原文
  • 融合注视点与轨迹建模时空上下文
  • 在两个数据集上达81.25%和81.85%的Dice分数
  • 适合医学影像分析与眼动研究交叉方向

医疗图像分割高度依赖耗时费力的像素级标注。眼动追踪提供了一种低成本且可自然融入临床流程的替代方案。眼动仪记录的注视信息通过固定点体现临床医生关注的空间区域,通过轨迹反映其逐步视觉感知的时间上下文。然而,有效建模时间轨迹仍具挑战,且探索性注视带来的噪声严重限制分割性能。为此,我们提出轨迹引导的不确定性感知网络(TrailNet),从空间语义建模逐步拓展到时间上下文建模,联合利用固定点与轨迹。具体而言,所提出的轨迹引导时空编码器建模时间上下文,并与图像空间语义建立互补交互以强化目标感知。此外,多尺度不确定性解码器利用类别互斥约束生成确定性预测,缓解由噪声引起的监督不确定性。为实现无眼动推理,我们进一步引入循环蒸馏策略,通过师生网络传递特征级知识。在两个公开数据集上的实验结果表明,TrailNet优于现有方法,分别取得81.25%和81.85%的Dice分数。

原文摘要 · Abstract (English)

Medical image segmentation heavily depends on labor-intensive and time-consuming pixel-level annotations. Eye tracking offers a cost-effective solution that can be naturally integrated into clinical workflows. Recorded by eye trackers, gaze conveys the spatial regions of clinicians' attention through fixations and the temporal context of clinicians' progressive visual perception from trajectories. Nevertheless, effective modeling of temporal trajectories remains challenging, and noise in gaze caused by exploratory fixations greatly limits segmentation performance. To overcome these limitations, we propose the Trajectory-guided Uncertainty-aware Network (TrailNet), which exploits gaze-supervised medical image segmentation from spatial semantics modeling to temporal context by jointly leveraging fixations and trajectories. Specifically, the proposed trajectory-guided spatio-temporal encoder models temporal context and establishes complementary interactions with image spatial semantics to strengthen target perception. Furthermore, the multi-scale uncertainty decoder leverages category mutual-exclusivity constraints to produce deterministic predictions and mitigate supervision uncertainty induced by noise. To enable gaze-free inference, we further introduce a cycle distillation strategy that transfers feature-level knowledge via teacher-student networks. Experimental results on two public datasets demonstrate that TrailNet outperforms state-of-the-art methods, achieving Dice scores of 81.25% and 81.85%, respectively.

医学图像弱监督眼动追踪分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。