用时空推理追踪手术注意力,提升机器人视野规划精度
SurgAtt-Tracker: Online Surgical Attention Tracking via Temporal Proposal Reranking and Motion-Aware Refinement
- 通过时序提议重排序与运动感知优化,非直接回归追踪注意力
- 在多数据集上达当前最优,对遮挡和多器械干扰鲁棒
- 适合做手术机器人视野自动控制的实时引导信号
精准稳定的视野(FoV)引导对微创手术的安全高效至关重要,但现有方法常将视觉注意力估计与相机控制混淆,或依赖对象中心假设。本文将手术注意力追踪建模为时空学习问题,以密集注意力热图形式表征术者关注区域,实现连续可解释的逐帧视野引导。提出SurgAtt-Tracker框架,通过提议级重排序与运动感知精修,利用时序一致性实现鲁棒追踪,而非直接回归。为支持系统训练与评估,构建了SurgAtt-1.16M大规模基准,采用临床合理标注协议,支持跨术式、跨机构的热图级注意力分析。在多个手术数据集上的实验证明,SurgAtt-Tracker在遮挡、多器械干扰及跨域设置下均表现优异,持续达到最先进水平。该方法不仅实现注意力追踪,还可输出逐帧视野引导信号,直接用于下游机器人视野规划与自动相机控制。
原文摘要 · Abstract (English)
Accurate and stable field-of-view (FoV) guidance is critical for safe and efficient minimally invasive surgery, yet existing approaches often conflate visual attention estimation with downstream camera control or rely on direct object-centric assumptions. In this work, we formulate surgical attention tracking as a spatio-temporal learning problem and model surgeon focus as a dense attention heatmap, enabling continuous and interpretable frame-wise FoV guidance. We propose SurgAtt-Tracker, a holistic framework that robustly tracks surgical attention by exploiting temporal coherence through proposal-level reranking and motion-aware refinement, rather than direct regression. To support systematic training and evaluation, we introduce SurgAtt-1.16M, a large-scale benchmark with a clinically grounded annotation protocol that enables comprehensive heatmap-based attention analysis across procedures and institutions. Extensive experiments on multiple surgical datasets demonstrate that SurgAtt-Tracker consistently achieves state-of-the-art performance and strong robustness under occlusion, multi-instrument interference, and cross-domain settings. Beyond attention tracking, our approach provides a frame-wise FoV guidance signal that can directly support downstream robotic FoV planning and automatic camera control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。