提出自适应视觉锚定,让多图问答更准更快
AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- 动态压缩无关视觉信息,保留关键图像特征
- 在多个数据集上提升准确率,最高达+6.2%
- 无需训练,适配各类大模型,适合多图问答场景
多模态大模型(MLLM)推动了视觉问答(VQA)从单图向多图发展。然而,多图输入带来的大量冗余视觉信息会降低回答准确率和效率。现有方法难以灵活控制压缩后的视觉标记数量,且常生成离散的视觉片段,影响模型对图像的整体理解。本文提出一种通用、无需训练的自适应视觉锚定策略,可无缝嵌入现有MLLM中,通过自适应压缩显著提升性能。同时引入协同解码机制,平衡全局与压缩视觉输入的效果,实现最优表现。大量实验验证了该方法的有效性,在多个主流模型和数据集上均取得一致提升,代码将公开。
原文摘要 · Abstract (English)
The advancement of Multimodal Large Language Models (MLLMs) has driven significant progress in Visual Question Answering (VQA), evolving from Single to Multi Image VQA (MVQA). However, the increased number of images in MVQA inevitably introduces substantial visual redundancy that is irrelevant to question answering, negatively impacting both accuracy and efficiency. To address this issue, existing methods lack flexibility in controlling the number of compressed visual tokens and tend to produce discrete visual fragments, which hinder MLLMs' ability to comprehend images holistically. In this paper, we propose a straightforward yet universal Adaptive Visual Anchoring strategy, which can be seamlessly integrated into existing MLLMs, offering significant accuracy improvements through adaptive compression. Meanwhile, to balance the results derived from both global and compressed visual input, we further introduce a novel collaborative decoding mechanism, enabling optimal performance. Extensive experiments validate the effectiveness of our method, demonstrating consistent performance improvements across various MLLMs. The code will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。