通过多模态大模型优化手术三元组识别中的任务冲突问题。
MEJO: MLLM-Engaged Surgical Triplet Recognition via Inter- and Intra-Task Joint Optimization
- 分离共享与特定任务表征,动态融合语义提示增强特征。
- 在CholecT45和CholecT50上准确率提升显著,解决长尾分布难题。
- 适合关注医疗视觉理解与多任务优化的研究者。
手术三元组识别需同时识别器械、动词、目标及其组合,面临长尾数据分布挑战。主流多任务学习虽具协同优势,但仍存在两大问题:任务间表征混杂导致优化冲突;类别不平衡引发任务内优化矛盾。为此,提出MLLM驱动的联合优化框架MEJO。针对任务间冲突,设计共享-特定-解耦(S²D)学习机制,将表征分解为共享与特定部分,并构建基于多模态大语言模型(MLLM)的动态概率提示池,以专家级语义线索增强视觉特征。同时,通过覆盖时空维度的任务专属提示建模,有效缓解任务间歧义。针对任务内冲突,提出协调梯度学习(CGL)策略,拆解并重平衡头尾类别的正负梯度,实现更协同的学习行为。在CholecT45和CholecT50数据集上的大量实验验证了该框架的有效性,显著提升了三元组识别性能。
原文摘要 · Abstract (English)
Surgical triplet recognition, which involves identifying instrument, verb, target, and their combinations, is a complex surgical scene understanding challenge plagued by long-tailed data distribution. The mainstream multi-task learning paradigm benefiting from cross-task collaborative promotion has shown promising performance in identifying triples, but two key challenges remain: 1) inter-task optimization conflicts caused by entangling task-generic and task-specific representations; 2) intra-task optimization conflicts due to class-imbalanced training data. To overcome these difficulties, we propose the MLLM-Engaged Joint Optimization (MEJO) framework that empowers both inter- and intra-task optimization for surgical triplet recognition. For inter-task optimization, we introduce the Shared-Specific-Disentangled (S$^2$D) learning scheme that decomposes representations into task-shared and task-specific components. To enhance task-shared representations, we construct a Multimodal Large Language Model (MLLM) powered probabilistic prompt pool to dynamically augment visual features with expert-level semantic cues. Additionally, comprehensive task-specific cues are modeled via distinct task prompts covering the temporal-spatial dimensions, effectively mitigating inter-task ambiguities. To tackle intra-task optimization conflicts, we develop a Coordinated Gradient Learning (CGL) strategy, which dissects and rebalances the positive-negative gradients originating from head and tail classes for more coordinated learning behaviors. Extensive experiments on the CholecT45 and CholecT50 datasets demonstrate the superiority of our proposed framework, validating its effectiveness in handling optimization conflicts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。