用眼神追踪技术分析手术室中医生注意力,提升对角色、阶段和沟通的理解。
Where are they looking in the operating room?

- 基于眼动热力图预测医生注视位置,无需额外标注
- 在两个数据集上实现角色识别F1 0.92,阶段识别F1 0.95
- 自监督建模眼动时空特征,显著提升团队沟通检测性能
目的:眼神追随(gaze-following)作为计算机视觉中推断个体注视方向的任务,已广泛应用于视觉注意力建模、社交场景理解与人机交互。然而,该任务尚未在手术室(OR)这一高风险、复杂环境中开展。本文首次将眼神追随引入外科领域,展示了其在理解临床角色、手术阶段与团队协作方面的巨大潜力。方法:我们扩展了4D-OR数据集以包含眼神追随标注,并在Team-OR数据集中新增眼神追随与团队沟通活动标注。提出新方法,利用眼神追随模型完成临床角色预测、手术阶段识别和团队沟通检测。角色与阶段识别采用仅依赖眼动预测的热力图方法;团队沟通检测则通过自监督方式训练时空模型,编码眼动片段特征,再输入时序活动检测模型。结果:在4D-OR与Team-OR数据集上的实验表明,本方法在所有下游任务中均达到当前最优表现。定量结果显示,角色预测F1得分为0.92,阶段识别达0.95。此外,在团队沟通检测任务中,性能较现有最佳基线提升超过30%。结论:本文首次将眼神追随引入手术室,开辟了外科数据科学的新方向,为计算机辅助手术中的手术流程分析提供了强大工具。
原文摘要 · Abstract (English)
Purpose: Gaze-following, the task of inferring where individuals are looking, has been widely studied in computer vision, advancing research in visual attention modeling, social scene understanding, and human-robot interaction. However, gaze-following has never been explored in the operating room (OR), a complex, high-stakes environment where visual attention plays an important role in surgical workflow analysis. In this work, we introduce the concept of gaze-following to the surgical domain, and demonstrate its great potential for understanding clinical roles, surgical phases, and team communications in the OR. Methods: We extend the 4D-OR dataset with gaze-following annotations, and extend the Team-OR dataset with gaze-following and a new team communication activity annotations. Then, we propose novel approaches to address clinical role prediction, surgical phase recognition, and team communication detection using a gaze-following model. For role and phase recognition, we propose a gaze heatmap-based approach that uses gaze predictions solely; for team communication detection, we train a spatial-temporal model in a self-supervised way that encodes gaze-based clip features, and then feed the features into a temporal activity detection model. Results: Experimental results on the 4D-OR and Team-OR datasets demonstrate that our approach achieves state-of-the-art performance on all downstream tasks. Quantitatively, our approach obtains F1 scores of 0.92 for clinical role prediction and 0.95 for surgical phase recognition. Furthermore, it significantly outperforms existing baselines in team communication detection, improving previous best performances by over 30%. Conclusion: We introduce gaze-following in the OR as a novel research direction in surgical data science, highlighting its great potential to advance surgical workflow analysis in computer-assisted interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。