诊断视觉推理缺陷,提出带轻量视觉模块的智能体架构提升准确性
Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward
- 用三阶段评估框架系统分析视觉语言模型失败模式
- 在MMMU和MathVista上分别提升10.3和6.0分,超越7B基线
- 适合关注视觉推理、多模态模型优化的研究者
整合视觉与文本推理的多模态大语言模型(MLLM)采用思维链(CoT)提示来处理复杂视觉任务,但仍存在视觉幻觉和过度依赖文本先验的问题。本文通过三阶段评估框架对前沿视觉语言模型进行系统诊断,揭示关键失效模式。为此,我们提出一种基于智能体的架构,将大语言模型推理与轻量级视觉模块结合,实现细粒度分析与推理链的迭代优化。实验结果表明,未来视觉推理模型应集成更广泛的专用工具以分析视觉内容。所提系统在MMMU上提升10.3分,在MathVista上提升6.0分,优于7B基线,媲美甚至超过更大模型。相关框架与评估套件将开源,推动后续研究。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) that integrate visual and textual reasoning leverage chain-of-thought (CoT) prompting to tackle complex visual tasks, yet continue to exhibit visual hallucinations and an over-reliance on textual priors. We present a systematic diagnosis of state-of-the-art vision-language models using a three-stage evaluation framework, uncovering key failure modes. To address these, we propose an agent-based architecture that combines LLM reasoning with lightweight visual modules, enabling fine-grained analysis and iterative refinement of reasoning chains. Our results highlight future visual reasoning models should focus on integrating a broader set of specialized tools for analyzing visual content. Our system achieves significant gains (+10.3 on MMMU, +6.0 on MathVista over a 7B baseline), matching or surpassing much larger models. We will release our framework and evaluation suite to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。