arXiv:2603.00108cs.ROcs.AI2026-03

提出多模态融合网络,提升真实手术中技能评估的准确性和可靠性。

SurgFusion-Net: Diversified Adaptive Multimodal Fusion Network for Surgical Skill Assessment

  • 设计自适应双注意力机制,动态融合视觉、运动与工具信息。
  • 在三个临床数据集上显著优于基线,最高提升0.0538(SCC)。
  • 首个面向真实手术场景的多模态数据集,适合临床研究与教育应用。

机器人辅助手术(RAS)已广泛应用于临床,利用多模态数据实现自动化手术技能评估具有变革潜力。然而,由于任务复杂、标注数据有限及跨模态融合技术不足,现有方法仍面临挑战。当前最优方案仅使用RGB视频且局限于模拟环境,难以应对真实手术中因器械运动、组织移动和视角变化带来的域差距。本文提出SurgFusion-Net与发散调节注意力(DRA)策略,首次构建两个临床级数据集:包含37例机器人辅助子宫切除术(RAH)的RAH-skill数据集(279,691帧),以及33例机器人辅助根治性前列腺切除术(RARP)的RARP-skill数据集(70,661帧),均含M-GEARS评分、光流和器械分割掩码。DRA通过自适应双注意力与多样性促进的多头注意力,基于手术上下文融合三模态信息,显著提升评估精度。在JIGSAWS、RAH-skill与RARP-skill上验证,该方法在留一患者/留一操作设置下,于不同任务中分别实现0.02、0.04的SCC提升,在RAH-skill与RARP-skill上分别取得0.0538与0.0493的绝对增益。

原文摘要 · Abstract (English)

Robotic-assisted surgery (RAS) is established in clinical practice, and automated surgical skill assessment utilizing multimodal data offers transformative potential for surgical analytics and education. However, developing effective multimodal methods remains challenging due to the task complexity, limited annotated datasets and insufficient techniques for cross-modal information fusion. Existing state-of-the-art relies exclusively on RGB video and only applies on dry-lab settings, failing to address the significant domain gap between controlled simulation and real clinical cases, where the surgical environment together with camera and tissue motion introduce substantial complexities. This work introduces SurgFusion-Net and Divergence Regulated Attention (DRA), an innovative fusion strategy for multimodal surgical skill assessment. We contribute two first-of-their-kind clinical datasets: the RAH-skill dataset containing 279,691 RGB frames from 37 videos of Robot-assisted Hysterectomy (RAH), and the RARP-skill dataset containing 70,661 RGB frames from 33 videos of Robot-Assisted Radical Prostatectomy (RARP). Both datasets include M-GEARS skill annotations, corresponding optical flow and tool segmentation masks. DRA incorporates adaptive dual attention and diversity-promoting multi-head attention to fuse multimodal information, from three modalities, based on surgical context, enhancing assessment accuracy and reliability. Validated on the JIGSAWS benchmark, RAH-skill, and RARP-skill datasets, our approach outperforms recent baselines with SCC improvements of 0.02 in LOSO, 0.04 in LOUO across JIGSAWS tasks, and 0.0538 and 0.0493 gains on RAH-skill and RARP-skill, respectively.

手术评估多模态融合临床数据集注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。