arXiv:2608.22972cs.CV2026-08中稿 · TMI

用专家混合机制融合手术知识,提升机器人手术动作识别准确率

Optimize Surgical Triplet Recognition: A Knowledge-Driven Mixture-of-Experts Solution

论文配图:Optimize Surgical Triplet Recognition: A Knowledge-Driven Mixture-of-Experts Solution
图 1 · 摘自论文原文
  • 分组件适配器解耦时空特征,缓解任务冲突
  • 自适应梯度重平衡策略提升罕见类别的识别性能
  • 融合大模型手术知识,增强模型可解释性与鲁棒性

手术动作三元组识别是语境感知机器人辅助手术中的关键任务,通过识别器械、动词、目标及其关联关系,实现自动手术行为理解。然而现有方法因三大问题难以分析复杂手术场景:(1) 特征空间纠缠导致组件级优化冲突;(2) 数据严重不均衡引发类别级优化冲突;(3) 缺乏领域知识引导,限制模型可解释性与鲁棒性。为此,我们提出基于知识驱动学习的专家混合协同优化框架(MoeCo)。在协同优化流程中,首先引入组件定制适配器,在时空域解耦任务特异性特征,促进组件专业化;其次设计协调梯度学习策略,自适应重平衡正负样本梯度,增强对稀有类别的感知能力;尤为关键的是,借鉴外科领域专业知识,提出知识驱动的专家混合机制,通过激活专家动态整合多模态大语言模型引导的知识,丰富协同优化过程中的表征能力。在公开数据集CholecT45和CholecT50上的大量实验验证了所提协同优化流程的有效性,以及知识驱动专家混合机制在动态先验融合方面的优势。

原文摘要 · Abstract (English)

Surgical action triplet recognition constitutes a critical task in context-aware robot-assisted surgery, facilitating automatic surgical action perception by identifying instrument, verb, target, and their association. However, existing works struggle to analyze such complex surgical scenes due to three main issues: (1) component-level optimization conflicts caused by entangled feature spaces, (2) category-level optimization conflicts arising from severe data imbalance, and (3) lack of domain knowledge guidance that limits model interpretability and robustness. To address these challenges, we propose a Mixture-of-Experts-guided Co-Optimization (\textit{MoeCo}) framework powered by knowledge-driven learning. Within the co-optimization pipeline, to first mitigate component-level conflicts, we introduce a component-tailored adapter that disentangles task-specific features across spatial-temporal regimes, facilitating effective component specialization. Next, we develop a coordinated gradient learning strategy to handle category-level conflicts, which adaptively rebalances positive-negative gradients to enhance the perception of rare categories. Notably, inspired by surgical domain expertise, we introduce a knowledge-driven mixture-of-experts mechanism that dynamically integrates multimodal large language model-guided knowledge via activated experts, thereby enriching the co-optimization pipeline with more expressive and robust representations. Extensive experiments on the public CholecT45 and CholecT50 datasets confirm the effectiveness of the proposed co-optimization pipeline and the superiority of dynamic priors integration via the knowledge-driven mixture-of-experts mechanism.

手术识别专家混合多模态知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。