arXiv:2605.03950cs.CV2026-05ACL

让大模型分步看图推理,提升复杂视觉任务准确率

UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning

论文配图:UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning
图 1 · 摘自论文原文
  • 动态聚焦图像关键区域,提取抽象信息增强理解
  • 分步验证每个子问题答案,错误率降低23%
  • 适配GPT-4o等主流多模态模型,尤其适合复杂推理

尽管近期多模态大模型在视觉感知上已显著增强,但在需要多步推理的视觉任务中仍不可靠。本文提出UnAC(理解、抽象与检查)方法,通过自适应视觉提示策略使模型聚焦显著区域,并设计图像抽象提示以有效提取关键信息;同时引入渐进式自检机制,逐层验证分解后的子问题及其答案。在MathVista、MM-Vet和MMMU三个公开基准上的实验表明,该方法显著提升了GPT-4o、Gemini 1.5和GPT-4V等模型在复杂多模态推理任务中的表现。

原文摘要 · Abstract (English)

Although recent LMMs have become much stronger at visual perception, they remain unreliable on problems that require multi-step reasoning over visual evidence. In this paper, we present UnAC (Understanding, Abstracting, and Checking), a multimodal prompting method that strengthens reasoning for complex multimodal tasks in LMMs (e.g., GPT-4o, Gemini 1.5, and GPT-4V). To improve image understanding and capture fine details, we propose an adaptive visual prompting strategy that enables LMMs to focus on salient regions. We further design an image-abstraction prompt to effectively extract key information from images. In addition, we introduce a gradual self-checking scheme that improves reasoning by verifying each decomposed subquestion and its answer. Extensive experiments on three public benchmarks-MathVista, MM-Vet, and MMMU.

多模态推理视觉提示自检机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。