arXiv:2511.03206cs.CVcs.AI2025-11EMNLP

提出QG-CoC框架,让多图理解更精准连贯

QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models

  • 用问题引导生成连贯的图像描述链,提升跨图感知
  • 在多图任务中显著优于现有提示方法,尤其在复杂场景
  • 零样本通用设计,适合各类多模态大模型使用

近期,多模态大语言模型(MLLMs)在多图像场景下面临两大挑战:(1) 对不同图像间细粒度视觉信息的感知不足;(2) 难以有效整合并推理多个视觉输入的信息。尽管已有多种提示方法用于描述视觉内容,但多数研究局限于单图像设置或特定受限场景,未能系统探讨MLLMs在更普遍、复杂的多图像推理任务中的表现。为此,我们首次系统性地调研了当前提示方法在多图像情境下对细粒度视觉细节的感知能力及信息处理机制。结果表明,现有方法在关注关键线索和融合感知与推理方面存在明显短板。基于此,我们提出一种新的零样本提示方法——问题引导的图像描述链(QG-CoC),该方法可灵活应对任意数量图像的任务。我们在多个开源与闭源的MLLM上评估该方法,在多图像与单图像基准测试中均表现出色,尤其在现有方法失效的挑战性场景中实现显著提升。

原文摘要 · Abstract (English)

Recently, Multimodal Large Language Models (MLLMs) encounter two key issues in multi-image contexts: (1) a lack of fine-grained perception across disparate images, and (2) a diminished capability to effectively reason over and synthesize information from multiple visual inputs. However, while various prompting methods aim to describe visual content, many existing studies focus primarily on single-image settings or specific, constrained scenarios. This leaves a critical gap in understanding and addressing how MLLMs tackle more general and complex multi-image reasoning tasks. Thus, we first extensively investigate how current prompting methods perceive fine-grained visual details and process visual information when dealing with multiple images. Our findings reveal that existing prompting methods fall short in attending to needed clues and seamlessly integrating perception and reasoning. Inspired by the findings, we propose a new zero-shot prompting method, Question-Guided Chain-of-Captions (QG-CoC), a generalized prompting approach that effectively handles problems with an arbitrary number of images. We evaluate our method on various open-source and closed-source MLLMs for multi-image and single-image benchmarks. Experimental results indicate that QG-CoC demonstrates competitive performance across tasks and exhibits robust improvements in the challenging scenarios where existing prompting methods fail.

多模态提示工程图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。