arXiv:2509.24776cs.CVcs.AI2025-09被引 3

让AI看图说话更靠谱:显式结合视觉与文本线索提升推理能力

VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding

  • 分两阶段训练:先强化感知,再用奖励机制引导推理
  • 在多个数据集上显著提升小模型的推理准确率和鲁棒性
  • 适合需要可解释、可靠多模态推理的应用场景

多模态大语言模型常难以基于感知证据进行推理。我们系统研究了四种感知策略(显式、隐式、视觉、文本)在四个多模态基准和两个MLLM上的表现。结果表明,显式感知,尤其是与文本提示结合时,能持续带来最佳提升,尤其对小模型效果显著。基于此,我们提出VTPerception-R1,一种统一的两阶段框架,将感知与推理解耦。第一阶段采用感知增强微调,第二阶段通过新颖的视觉、文本和一致性奖励,实施感知感知强化学习。实验表明,该方法显著提升了多样化任务中的推理准确率与鲁棒性,提供了一种可扩展且可审计的感知基多模态推理方案。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) often struggle to ground reasoning in perceptual evidence. We present a systematic study of perception strategies-explicit, implicit, visual, and textual-across four multimodal benchmarks and two MLLMs. Our findings show that explicit perception, especially when paired with textual cues, consistently yields the best improvements, particularly for smaller models. Based on this insight, we propose VTPerception-R1, a unified two-stage framework that decouples perception from reasoning. Stage 1 introduces perception-augmented fine-tuning, and Stage 2 applies perception-aware reinforcement learning with novel visual, textual, and consistency rewards. Experiments demonstrate that VTPerception-R1 significantly improves reasoning accuracy and robustness across diverse tasks, offering a scalable and auditable solution for perception-grounded multimodal reasoning. Our code is available at: https://github.com/yizhuoDi/VTPerceprion-R1.

多模态推理视觉理解强化学习模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。