arXiv:2509.16618cs.CVcs.AI2025-09中稿 · MICCAI2025被引 5

用Mamba2增强视觉问答,让机器人手术模型更懂空间细节。

Surgical-MambaLLM: Mamba2-enhanced Multimodal Large Language Model for VQLA in Robotic Surgery

  • 用双向Mamba2融合图像与文本,捕捉跨模态依赖关系。
  • 在EndoVis17/18数据集上超越现有方法,提升显著。
  • 专为手术器械布局设计扫描模式,强化空间理解能力。

近年来,机器人手术中的视觉问题定位回答(Surgical-VQLA)受到关注,有助于医学生和初级医生理解手术场景。尽管大型语言模型(LLMs)的发展提供了新思路,但现有方法难以建立文本与视觉细节间的复杂关联,且对手术场景的空间信息感知不足。为此,我们提出Surgical-MambaLLM,首次将Mamba2与LLM结合应用于手术领域,利用Mamba2有效捕捉跨模态依赖并感知手术场景的空间信息,从而提升LLM对手术图像的理解。具体地,提出跨模态双向Mamba2融合(CBMI)模块,实现高效多模态融合;针对手术场景的几何特性,设计外科器械感知(SIP)扫描模式,使Mamba2更精准扫描手术图像,增强空间理解。大量实验表明,Surgical-MambaLLM在EndoVis17-VQLA和EndoVis18-VQLA数据集上均优于当前最优方法,显著提升Surgical-VQLA任务性能。

原文摘要 · Abstract (English)

In recent years, Visual Question Localized-Answering in robotic surgery (Surgical-VQLA) has gained significant attention for its potential to assist medical students and junior doctors in understanding surgical scenes. Recently, the rapid development of Large Language Models (LLMs) has provided more promising solutions for this task. However, current methods struggle to establish complex dependencies between text and visual details, and have difficulty perceiving the spatial information of surgical scenes. To address these challenges, we propose a novel method, Surgical-MambaLLM, which is the first to combine Mamba2 with LLM in the surgical domain, that leverages Mamba2's ability to effectively capture cross-modal dependencies and perceive spatial information in surgical scenes, thereby enhancing the LLMs' understanding of surgical images. Specifically, we propose the Cross-modal Bidirectional Mamba2 Integration (CBMI) module to leverage Mamba2 for effective multimodal fusion, with its cross-modal integration capabilities. Additionally, tailored to the geometric characteristics of surgical scenes, we design the Surgical Instrument Perception (SIP) scanning mode for Mamba2 to scan the surgical images, enhancing the model's spatial understanding of the surgical scene. Extensive experiments demonstrate that our Surgical-MambaLLM model outperforms the state-of-the-art methods on the EndoVis17-VQLA and EndoVis18-VQLA datasets, significantly improving the performance of the Surgical-VQLA task.

手术视觉多模态Mamba2问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。