arXiv:2602.09259cs.ROcs.HC2026-02

用虚拟手术模拟数据对比主动与被动注视,发现被动注视可替代部分专家注视用于训练模型。

Data-centric Design of Learning-based Surgical Gaze Perception Models in Multi-Task Simulation

  • 构建了主动执行与被动观看双模式手术注视数据集,实现同一视频下对比分析。
  • 被动注视可恢复中间水平主动注视的大部分注意力模式,但精度有降低。
  • 新手的被动注视标签对高质量演示效果影响小,适合规模化众包标注。

在机器人辅助微创手术中,触觉反馈和深度线索减少,使专家视觉感知变得关键,推动了基于注视的训练和学习型手术感知模型的发展。然而,操作专家注视数据收集成本高,且尚不清楚注视监督来源(技能水平:中级/初级,感知模态:主动执行/被动观察)如何影响注意力模型的学习效果。本研究在da Vinci SimNow模拟器上,针对四项任务构建了配对的主动-被动多任务手术注视数据集。主动注视通过佩戴带眼动追踪的VR头显在任务执行时记录,对应的视频作为刺激材料,用于收集观察者的被动注视,实现同一视频下的可控对比。我们量化了技能与模态依赖的注视组织差异,并通过凝视密度重叠分析和单帧显著性建模评估被动注视替代操作监督的可行性。结果表明,MSI-Net在所有设置下均产生稳定且可解释的预测,而SalGAN则表现不稳定且常与人类注视不一致。以被动注视训练的模型可恢复大量中级主动注意力,但存在可预测的退化,且主动与被动之间的迁移呈非对称性。值得注意的是,初级观察者的被动标签在高质量示范下仅带来有限损失,表明其为手术指导与感知建模提供了一条可行的可扩展众包标注路径。

原文摘要 · Abstract (English)

In robot-assisted minimally invasive surgery (RMIS), reduced haptic feedback and depth cues increase reliance on expert visual perception, motivating gaze-guided training and learning-based surgical perception models. However, operative expert gaze is costly to collect, and it remains unclear how the source of gaze supervision, both expertise level (intermediate vs. novice) and perceptual modality (active execution vs. passive viewing), shapes what attention models learn. We introduce a paired active-passive, multi-task surgical gaze dataset collected on the da Vinci SimNow simulator across four drills. Active gaze was recorded during task execution using a VR headset with eye tracking, and the corresponding videos were reused as stimuli to collect passive gaze from observers, enabling controlled same-video comparisons. We quantify skill- and modality-dependent differences in gaze organization and evaluate the substitutability of passive gaze for operative supervision using fixation density overlap analyses and single-frame saliency modeling. Across settings, MSI-Net produced stable, interpretable predictions, whereas SalGAN was unstable and often poorly aligned with human fixations. Models trained on passive gaze recovered a substantial portion of intermediate active attention, but with predictable degradation, and transfer was asymmetric between active and passive targets. Notably, novice passive labels approximated intermediate-passive targets with limited loss on higher-quality demonstrations, suggesting a practical path for scalable, crowd-sourced gaze supervision in surgical coaching and perception modeling.

手术视觉注视预测数据设计众包标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。