arXiv:2508.00356cs.CVcs.MA2025-08

用双智能体协作框架,让大模型更准地分析多张图片。

Analyze-Prompt-Reason: A Collaborative Agent-Based Framework for Multi-Image Vision-Language Reasoning

  • 设计语言提示生成器与视觉推理模型协同工作
  • 在18个数据集上实现接近顶尖的多图理解性能
  • 无需训练,适合各类图像推理任务

我们提出一种基于智能体的协作框架,用于多图像视觉语言推理。该方法通过双智能体系统应对跨多样数据集和任务格式的交错多模态推理挑战:语言型提示工程师生成上下文感知、任务特定的提示,视觉推理器(大视觉语言模型,LVLM)负责最终推理。框架完全自动化、模块化且无需训练,可泛化至分类、问答及自由生成等涉及一张或多张输入图像的任务。我们在2025 MIRAGE Challenge(Track A)的18个多样化数据集上评估该方法,涵盖文档问答、视觉对比、对话理解及场景级推理等多种视觉推理任务。结果表明,当获得有效提示引导时,LVLM能高效处理多图像推理。值得注意的是,Claude 3.7在复杂任务上表现接近极限:TQA准确率达99.13%,DocVQA达96.87%,MMCoQA ROUGE-L为75.28。我们还研究了模型选择、样本数量和输入长度等设计因素对不同LVLM推理性能的影响。

原文摘要 · Abstract (English)

We present a Collaborative Agent-Based Framework for Multi-Image Reasoning. Our approach tackles the challenge of interleaved multimodal reasoning across diverse datasets and task formats by employing a dual-agent system: a language-based PromptEngineer, which generates context-aware, task-specific prompts, and a VisionReasoner, a large vision-language model (LVLM) responsible for final inference. The framework is fully automated, modular, and training-free, enabling generalization across classification, question answering, and free-form generation tasks involving one or multiple input images. We evaluate our method on 18 diverse datasets from the 2025 MIRAGE Challenge (Track A), covering a broad spectrum of visual reasoning tasks including document QA, visual comparison, dialogue-based understanding, and scene-level inference. Our results demonstrate that LVLMs can effectively reason over multiple images when guided by informative prompts. Notably, Claude 3.7 achieves near-ceiling performance on challenging tasks such as TQA (99.13% accuracy), DocVQA (96.87%), and MMCoQA (75.28 ROUGE-L). We also explore how design choices-such as model selection, shot count, and input length-influence the reasoning performance of different LVLMs.

多图推理智能体协作大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。