提出手术动作三元组实例分割新任务,实现精准空间定位。
Grounding Surgical Action Triplets with Instrument Instance Segmentation: A Dataset and Target-Aware Fusion Approach
- 用实例分割关联器械、动作与目标,实现空间定位
- 构建3万+帧的标注数据集CholecTriplet-Seg,支持实例级评估
- 提出目标感知融合网络,提升解剖目标预测精度
理解手术器械-组织交互不仅需要识别器械执行的动作及其作用目标,还需在手术场景中精确定位这些交互。现有方法多依赖帧级分类,无法可靠关联动作与具体器械实例。此前的空间定位尝试主要基于类别激活图,精度和鲁棒性不足。为此,本文提出将手术动作三元组与器械实例分割结合的新任务——三元组分割(triplet segmentation),生成具有空间位置的<器械, 动作, 目标>输出。首先构建了包含超过30,000帧标注的大型数据集CholecTriplet-Seg,首次建立强监督、实例级别的三元组定位基准。为学习三元组分割,提出TargetFusionNet,一种在Mask2Former基础上引入目标感知融合机制的新架构,通过融合弱解剖先验与器械实例查询,解决解剖目标预测难题。在识别、检测及三元组分割指标上,TargetFusionNet持续优于现有基线,证明强实例监督与弱目标先验结合能显著提升手术动作理解的准确性和鲁棒性。三元组分割为手术场景理解提供统一框架,推动更可解释的智能手术分析。
原文摘要 · Abstract (English)
Understanding surgical instrument-tissue interactions requires not only identifying which instrument performs which action on which anatomical target, but also grounding these interactions spatially within the surgical scene. Existing surgical action triplet recognition methods are limited to learning from frame-level classification, failing to reliably link actions to specific instrument instances.Previous attempts at spatial grounding have primarily relied on class activation maps, which lack the precision and robustness required for detailed instrument-tissue interaction analysis.To address this gap, we propose grounding surgical action triplets with instrument instance segmentation, or triplet segmentation for short, a new unified task which produces spatially grounded <instrument, verb, target> outputs.We start by presenting CholecTriplet-Seg, a large-scale dataset containing over 30,000 annotated frames, linking instrument instance masks with action verb and anatomical target annotations, and establishing the first benchmark for strongly supervised, instance-level triplet grounding and evaluation.To learn triplet segmentation, we propose TargetFusionNet, a novel architecture that extends Mask2Former with a target-aware fusion mechanism to address the challenge of accurate anatomical target prediction by fusing weak anatomy priors with instrument instance queries.Evaluated across recognition, detection, and triplet segmentation metrics, TargetFusionNet consistently improves performance over existing baselines, demonstrating that strong instance supervision combined with weak target priors significantly enhances the accuracy and robustness of surgical action understanding.Triplet segmentation establishes a unified framework for spatially grounding surgical action triplets. The proposed benchmark and architecture pave the way for more interpretable, surgical scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。