用专家混合机制融合手术知识,提升机器人手术动作识别准确率
Optimize Surgical Triplet Recognition: A Knowledge-Driven Mixture-of-Experts Solution

- 分组件适配器解耦时空特征,缓解任务冲突
- 自适应梯度重平衡策略提升罕见类别的识别性能
- 融合大模型手术知识,增强模型可解释性与鲁棒性
手术动作三元组识别是语境感知机器人辅助手术中的关键任务,通过识别器械、动词、目标及其关联关系,实现自动手术行为理解。然而现有方法因三大问题难以分析复杂手术场景:(1) 特征空间纠缠导致组件级优化冲突;(2) 数据严重不均衡引发类别级优化冲突;(3) 缺乏领域知识引导,限制模型可解释性与鲁棒性。为此,我们提出基于知识驱动学习的专家混合协同优化框架(MoeCo)。在协同优化流程中,首先引入组件定制适配器,在时空域解耦任务特异性特征,促进组件专业化;其次设计协调梯度学习策略,自适应重平衡正负样本梯度,增强对稀有类别的感知能力;尤为关键的是,借鉴外科领域专业知识,提出知识驱动的专家混合机制,通过激活专家动态整合多模态大语言模型引导的知识,丰富协同优化过程中的表征能力。在公开数据集CholecT45和CholecT50上的大量实验验证了所提协同优化流程的有效性,以及知识驱动专家混合机制在动态先验融合方面的优势。
原文摘要 · Abstract (English)
Surgical action triplet recognition constitutes a critical task in context-aware robot-assisted surgery, facilitating automatic surgical action perception by identifying instrument, verb, target, and their association. However, existing works struggle to analyze such complex surgical scenes due to three main issues: (1) component-level optimization conflicts caused by entangled feature spaces, (2) category-level optimization conflicts arising from severe data imbalance, and (3) lack of domain knowledge guidance that limits model interpretability and robustness. To address these challenges, we propose a Mixture-of-Experts-guided Co-Optimization (\textit{MoeCo}) framework powered by knowledge-driven learning. Within the co-optimization pipeline, to first mitigate component-level conflicts, we introduce a component-tailored adapter that disentangles task-specific features across spatial-temporal regimes, facilitating effective component specialization. Next, we develop a coordinated gradient learning strategy to handle category-level conflicts, which adaptively rebalances positive-negative gradients to enhance the perception of rare categories. Notably, inspired by surgical domain expertise, we introduce a knowledge-driven mixture-of-experts mechanism that dynamically integrates multimodal large language model-guided knowledge via activated experts, thereby enriching the co-optimization pipeline with more expressive and robust representations. Extensive experiments on the public CholecT45 and CholecT50 datasets confirm the effectiveness of the proposed co-optimization pipeline and the superiority of dynamic priors integration via the knowledge-driven mixture-of-experts mechanism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。