让AI先看懂图再回答,提升视觉推理准确性
See First, Answer Later: Visual Evidence Pre-Alignment via Sufficiency-Driven RL

- 用足够性驱动强化学习优化视觉描述
- 在多个评测中显著提升视觉任务表现
- 适合需要强视觉理解的模型优化场景
多模态大语言模型虽融合文本推理与视觉输入,但响应常与图像不一致,表明推理时未能有效利用视觉证据。现有训练范式依赖大规模图文标题预训练,再经监督微调和强化学习实现指令遵循与复杂推理,但预训练提供的视觉对齐较弱:短而粗略的标题使模型偏向显著对象,忽略细粒度视觉信息。本文提出视觉证据预对齐(VEPA),作为预训练与后训练之间的中间阶段,采用组相对策略优化(GRPO)实现基于问题的视觉证据描述优化。大量实验表明,VEPA在多样基准上持续提升视觉密集型任务表现,并可补充标准监督后训练。进一步分析显示,性能提升源于更强且可迁移的视觉对齐能力,而非额外的任务特定训练。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) integrate strong text reasoning with visual inputs, yet their responses can be inconsistent with the underlying images, indicating ineffective utilization of visual evidence during inference. The prevailing training paradigm relies on large-scale caption-based pretraining for general alignment, followed by supervised fine-tuning and reinforcement learning to enable instruction following and complex reasoning. However, such pretraining provides only weak visual grounding: short, coarse captions bias models toward salient objects while neglecting fine-grained visual evidence. In this paper, we introduce Visual Evidence Pre-Alignment (VEPA), an intermediate stage between pretraining and post-training that explores a novel sufficiency-driven objective with Group Relative Policy Optimization (GRPO) to optimize question-conditioned visual evidence descriptions. Extensive experiments across diverse benchmarks show that our VEPA consistently enhances performance on visually demanding evaluations and complements standard supervised post-training. Further analyses show that the income stems from strengthened, transferable visual grounding, rather than from additional task-specific training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。