arXiv:2508.02886cs.CL2025-08被引 1

让视觉语言模型像人一样一步步推理,自检纠错提升复杂问题理解力。

Coherent Multimodal Reasoning with Iterative Self-Evaluation for Vision-Language Models

  • 分步拆解问题,逐轮自我评估修正推理路径。
  • 在多个基准上达69.4%平均准确率,领先开源模型2.4个百分点。
  • 适合需要深度跨模态推理的场景,如智能助手、自动驾驶理解。

尽管取得显著进展,当前大型语言模型(LLMs)和视觉语言模型(LVLMs)在处理复杂、多步骤的跨模态常识推理任务时仍存在不足,常表现为缺乏‘深思熟虑’,依赖表面关联而非深层链式推理,尤其在整合视觉信息与抽象概念时。为此,我们提出一致的多模态推理框架(CMRF),通过迭代自评估机制增强LVLM的常识推理能力。CMRF模仿人类解题方式,将复杂问题分解为子问题,生成逐步推理,并自我纠正错误。框架包含三个核心模块:推理分解单元(RDU)、上下文推理引擎(CIE)和一致性评估模块(CAM),结合自适应迭代精炼策略,系统性优化推理路径。基于LLaVA-1.6-34B,在新构建的多模态日常活动推理(MDAR)数据集上训练,CMRF在VCR、A-OKVQA和DailyLife-MRC等挑战性基准上达到开源模型最优表现,平均准确率达69.4%,较最佳基线提升2.4个百分点,尤其在复杂推理场景中优势明显。大量消融实验与人工评估证实各模块的关键作用及迭代精炼的有效性。

原文摘要 · Abstract (English)

Despite significant advancements, current large language models (LLMs) and vision-language models (LVLMs) continue to struggle with complex, multi-step, cross-modal common sense reasoning tasks, often exhibiting a lack of "deliberative thinking." They tend to rely on superficial associations rather than deep, chained inference, particularly when integrating visual information with abstract concepts. To address this, we propose the Coherent Multimodal Reasoning Framework (CMRF), a novel approach that enhances LVLMs' common sense reasoning capabilities through an iterative, self-evaluating inference mechanism. CMRF mimics human problem-solving by decomposing complex queries, generating step-by-step inferences, and self-correcting errors. Our framework integrates three key modules: a Reasoning Decomposition Unit (RDU) for breaking down problems into sub-questions, a Contextual Inference Engine (CIE) for contextual inference, and a Coherence Assessment Module (CAM) for evaluating logical consistency and confidence. Coupled with an Adaptive Iterative Refinement strategy, CMRF systematically refines its reasoning paths. Built upon LLaVA-1.6-34B and trained on a novel Multimodal Daily Activity Reasoning (MDAR) dataset, CMRF achieves state-of-the-art performance among open-source LVLMs on challenging benchmarks like VCR, A-OKVQA, and DailyLife-MRC. It attains an average accuracy of 69.4%, surpassing the best open-source baseline by +2.4 percentage points, with particular strength in complex reasoning scenarios. Extensive ablation studies and human evaluations confirm the critical contributions of each module and the effectiveness of iterative refinement in fostering more coherent and accurate reasoning.

多模态推理自评估视觉语言模型链式推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。