arXiv:2604.09051cs.CVcs.RO2026-04

精准分割机器人肾部分切除术中的缝合动作,提升手术智能分析能力

Fine-Grained Action Segmentation for Renorrhaphy in Robot-Assisted Partial Nephrectomy

  • 基于I3D特征构建时序模型,识别视觉相似但持续时间各异的缝合动作
  • DiffAct在多指标中表现最优,帧级准确率和编辑分数均领先
  • 适用于微创手术智能辅助,尤其适合需要精细动作识别的场景

机器人辅助部分肾切除术中肾缝合阶段的细粒度动作分割,面临帧级识别视觉相似缝合动作、持续时间不一及类别严重失衡的挑战。SIA-RAPN基准在50段临床视频上定义该问题,使用da Vinci Xi系统采集并标注12类帧级标签。比较了基于I3D特征的四种时序模型:MS-TCN++、AsFormer、TUT和DiffAct。评估指标包括平衡准确率、编辑分数、重叠阈值为10、25和50的分段F1、帧级准确率和帧级平均精度。除在五个已发布划分配置上的主评估外,还报告了在独立单孔肾部分切除术数据集上的跨域结果。在五次运行中最强报告值下,DiffAct在分段F1、帧级准确率、编辑分数和帧级mAP上均居首位,而MS-TCN++在平衡准确率上最高。

原文摘要 · Abstract (English)

Fine-grained action segmentation during renorrhaphy in robot-assisted partial nephrectomy requires frame-level recognition of visually similar suturing gestures with variable duration and substantial class imbalance. The SIA-RAPN benchmark defines this problem on 50 clinical videos acquired with the da Vinci Xi system and annotated with 12 frame-level labels. The benchmark compares four temporal models built on I3D features: MS-TCN++, AsFormer, TUT, and DiffAct. Evaluation uses balanced accuracy, edit score, segmental F1 at overlap thresholds of 10, 25, and 50, frame-wise accuracy, and frame-wise mean average precision. In addition to the primary evaluation across five released split configurations on SIA-RAPN, the benchmark reports cross-domain results on a separate single-port RAPN dataset. Across the strongest reported values over those five runs on the primary dataset, DiffAct achieves the highest F1, frame-wise accuracy, edit score, and frame mAP, while MS-TCN++ attains the highest balanced accuracy.

手术分割动作识别机器人手术时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。