用多模态融合预测手术可操作区域,助力机器人自主手术
SurgAM: Surgical Affordance Map Prediction with Multimodal Feature Fusion for Robot Autonomy

- 结合自监督视觉变压器与生成模型,融合语义与空间信息
- 在三种基础手术动作上实现当前最优性能,准确率显著提升
- 适合机器人手术自动化研究者,尤其关注场景理解与决策
手术自动化日益受到关注,但如何将视觉场景理解与自主动作规划相连接仍是核心挑战。尽管已有大量研究聚焦于场景感知(如器械识别与场景分割),对可执行操作的识别与预测仍缺乏深入探索。本文提出手术可操作性预测,旨在从视觉数据中识别基础手术动作的可行区域。为此,我们设计了一种自适应特征融合框架,利用自监督视觉变压器在语义理解上的优势,以及大规模生成模型在空间感知上的能力。进一步引入分层提示学习机制以适应不同手术流程上下文,并提出场景引导注意力解码器,聚焦关键手术区域并抑制背景干扰。为验证有效性,我们构建了一个新数据集,基于公开手术数据集标注了三种基础手术动作(抽吸、夹闭、牵拉)的可操作性。大量实验表明,本方法达到当前最优性能。此外,在肺和前列腺模拟体上验证了其在下游自动化的适用性,结果表明预测的可操作图能成功支持自主手术操作。
原文摘要 · Abstract (English)
Surgical automation is being increasingly studied, yet bridging visual scene understanding with autonomous action planning remains a fundamental challenge. While much research effort has been made on scene perception (e.g., tool recognition and scene segmentation), understanding and predicting actionable possibilities for surgical automation is still underexplored. In this paper, we introduce surgical affordance prediction, which identifies actionable regions for fundamental surgical actions from visual data. Specifically, a novel adaptive feature fusion framework is proposed that leverages the complementary strengths of a self-supervised vision transformer encoder for its superior semantic understanding and a large-scale generative model encoder for its spatially-aware capability. Furthermore, we introduce a hierarchical prompt learning mechanism to adapt to varying procedural contexts. Finally, a scene-guided attention decoder is proposed to focus on critical surgical areas while suppressing background distractions. To validate the effectiveness, we established a new dataset, derived from publicly available surgical datasets with affordance annotations for three basic surgical actions: aspiration, clipping, and retraction. Extensive experiments demonstrate that our approach achieves state-of-the-art performance. Moreover, we validate our framework's applicability for downstream automation on a realistic lung and prostate phantom, and results show that the predicted affordance maps successfully enable autonomous surgical actions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。