arXiv:2411.10937cs.CVcs.CL2024-11被引 6

用自洽提问增强记忆,提升手术视觉问答理解能力

Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry

  • 通过自生成记忆弥补多对象推理缺陷
  • 在三个数据集上达到最新最好效果
  • 适合需要强场景理解的医疗AI研究者

手术视觉问答(Surgical VQA)需对多个物体进行推理以全面理解手术场景。现有方法依赖跨模态融合,但常受限于场景理解不足和问题解析能力弱,且部分方法依赖外部资源(如预提取物体特征),易引入误差且泛化性差。为此,我们提出SCAN框架,一种基于多模态大模型的记忆增强方法,通过自洽提问实现上下文增强。SCAN自主生成两类记忆:直接记忆(DM)提供答案候选或提示,间接记忆(IM)由自洽的问题-提示对构成,用于捕捉更广泛的场景上下文。通过对这些面向对象的记忆进行推理,模型能更准确地理解图像并回答问题。在三个公开手术VQA数据集上的大量实验表明,SCAN达到当前最优性能,显著提升了各类手术场景下的准确率与鲁棒性。

原文摘要 · Abstract (English)

Comprehensively understanding surgical scenes in Surgical Visual Question Answering (Surgical VQA) requires reasoning over multiple objects. Previous approaches address this task using cross-modal fusion strategies to enhance reasoning ability. However, these methods often struggle with limited scene understanding and question comprehension, and some rely on external resources (e.g., pre-extracted object features), which can introduce errors and generalize poorly across diverse surgical environments. To address these challenges, we propose SCAN, a simple yet effective memory-augmented framework that leverages Multimodal LLMs to improve surgical context comprehension via Self-Contained Inquiry. SCAN operates autonomously, generating two types of memory for context augmentation: Direct Memory (DM), which provides multiple candidates (or hints) to the final answer, and Indirect Memory (IM), which consists of self-contained question-hint pairs to capture broader scene context. DM directly assists in answering the question, while IM enhances understanding of the surgical scene beyond the immediate query. Reasoning over these object-aware memories enables the model to accurately interpret images and respond to questions. Extensive experiments on three publicly available Surgical VQA datasets demonstrate that SCAN achieves state-of-the-art performance, offering improved accuracy and robustness across various surgical scenarios.

手术视觉问答多模态大模型记忆增强自洽提问

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。