arXiv:2607.21722cs.CV2026-07被引 1

用一致性约束提升视觉语言模型的推理可靠性。

Be Consistent! Enhancing Robust Visual Reasoning in LVLMs with Consistency Constraints

论文配图:Be Consistent! Enhancing Robust Visual Reasoning in LVLMs with Consistency Constraints
图 1 · 摘自论文原文
  • 设计新基准ConVBench,含六类复杂视觉推理任务。
  • 引入逻辑一致性和鲁棒准确率双指标,量化模型表现。
  • 提出ConVLM框架,通过奖励机制让模型对等价问题答一致。

尽管大型视觉语言模型具备强大的感知能力,但在视觉推理任务中仍显脆弱。现有评测基准多聚焦符号化数学或科学问题及简单视觉任务,难以评估复杂视觉推理与逻辑一致性——这正是可靠推理系统的关键要求。我们提出ConVBench,一个复杂的视觉中心推理基准,每个图像配对两个逻辑等价的问题,涵盖六类:动作与状态、复杂计数、空间推理、因果与意图理解、常识推理和时间感知。为配合该基准,我们定义了逻辑一致性与鲁棒准确率两项评估指标,联合衡量模型响应的正确性与一致性。我们进一步提出ConVLM,基于组相对策略优化(GRPO)的强化学习方法,引入新颖的一致性奖励。该方法利用自动生成的逻辑等价问题-答案对,采用准确性与一致性双信号奖励设计,促使模型在成对问题上保持回答一致。该框架在有或无严格答案监督下均有效。

原文摘要 · Abstract (English)

While Large Vision-Language Models (LVLMs) exhibit strong perceptual capabilities, they remain vulnerable in visual reasoning tasks. Existing benchmarks largely focus on symbolic mathematical or scientific problems and simple vision-centric tasks, offering limited assessment of complex visual reasoning and logical consistency, a critical requirement for reliable reasoning systems. We introduce ConVBench, a complex vision-centric reasoning benchmark in which each image is paired with two logically equivalent questions across six categories: action and state, complex counting, spatial reasoning, causal and intent understanding, commonsense reasoning, and temporal perception. To complement this benchmark, we define two evaluation metrics, logical consistency and robust accuracy, that jointly assess both the correctness and consistency of model responses. We further present ConVLM, which improves LVLM reasoning through Group Relative Policy Optimization (GRPO)-based reinforcement learning with a novel consistency reward. This method leverages automatically generated logically equivalent question-answer pairs and a dual-reward design combining accuracy- and consistency-based signals, encouraging agreement between paired responses. The framework functions effectively with or without strict answer supervision.

视觉推理一致性强化学习基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。