arXiv:2512.12487cs.CV2025-12被引 1

提升视觉语言模型的视觉提取与逻辑一致性

More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models

  • 分步优化视觉感知与文本推理,避免两者互相干扰
  • 在多个基准上将准确率从63.3%提升至68.8%
  • 适合需要可靠多模态推理的应用场景

强化学习通过可验证奖励(RLVR)已被扩展至视觉语言模型(VLMs),以激发长链多模态推理。然而,经过RLVR训练的VLMs仍存在两个顽固缺陷:视觉提取不准确(遗漏或虚构细节)和推理链条逻辑不一致,主要因可验证信号仅监督最终答案。本文提出PeRL-VL(视觉语言模型的感知与推理学习)框架,将视觉感知与文本推理分离优化。在感知方面,引入基于VLM的描述奖励,评估模型自生成图像描述的忠实性与充分性;在推理方面,增加仅文本的推理SFT阶段,利用逻辑丰富的思维链数据独立提升连贯性与逻辑一致性。在多个多模态基准上,PeRL-VL将平均Pass@1准确率从基础模型Qwen2.5-VL-7B的63.3%提升至68.8%,优于标准RLVR、纯文本推理SFT及直接从GPT-4o蒸馏的多模态方法。

原文摘要 · Abstract (English)

Reinforcement learning from verifiable rewards (RLVR) has recently been extended from text-only LLMs to vision-language models (VLMs) to elicit long-chain multimodal reasoning. However, RLVR-trained VLMs still exhibit two persistent failure modes: inaccurate visual extraction (missing or hallucinating details) and logically inconsistent chains-of-thought, largely because verifiable signals supervise only the final answer. We propose PeRL-VL (Perception and Reasoning Learning for Vision-Language Models), a decoupled framework that separately improves visual perception and textual reasoning on top of RLVR. For perception, PeRL-VL introduces a VLM-based description reward that scores the model's self-generated image descriptions for faithfulness and sufficiency. For reasoning, PeRL-VL adds a text-only Reasoning SFT stage on logic-rich chain-of-thought data, enhancing coherence and logical consistency independently of vision. Across diverse multimodal benchmarks, PeRL-VL improves average Pass@1 accuracy from 63.3% (base Qwen2.5-VL-7B) to 68.8%, outperforming standard RLVR, text-only reasoning SFT, and naive multimodal distillation from GPT-4o.

视觉语言模型推理优化多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。