arXiv:2604.13993cs.AIcs.CL2026-04

通过设计不同奖励信号,提升视觉语言模型的物理推理能力。

Reward Design for Physical Reasoning in Vision-Language Models

论文配图:Reward Design for Physical Reasoning in Vision-Language Models
图 1 · 摘自论文原文
  • 用四种奖励信号对比实验,探索如何优化模型推理行为。
  • 准确率奖励带来整体最好效果,注意力奖励显著提升空间推理。
  • 无需人工标注空间信息,模型自注意力可作有效监督信号。

视觉输入中的物理推理需要视觉感知、领域知识与多步符号推理的紧密结合。尽管当前最先进的视觉语言模型(VLMs)在物理基准测试中仍远低于人类表现,但后训练算法如监督微调(SFT)和组相对策略优化(GRPO)已在语言模型中展现强大推理提升。然而,奖励设计如何影响VLM的物理推理行为尚不明确。本文针对GRPO-based VLM训练开展系统性奖励消融研究,比较四种语义丰富度递增的奖励信号:格式合规性、答案准确性、综合评分奖励(答案正确性、物理原理识别、单位一致性)以及一种基于模型注意力权重的新型内部奖励。在包含3,000个问题的PhyX基准上评估,覆盖六个物理领域与六种推理类型,采用IBM Granite Vision 3.3(2B)。结果显示,在两种题型下,基于准确率的奖励使GRPO优于SFT,但提升幅度因奖励类型与领域而异。奖励设计并非统一提升性能,而是诱导特定领域的推理模式:准确率奖励获得最强整体提升;评分奖励改善结构化推理质量,但未带来一致准确率提升;注意力奖励增强空间推理,却降低符号域表现。内部注意力奖励无需空间标注,将空间关系准确率从0.27提升至0.50,表明监督模型生成时的关注区域是视觉引导物理推理的有前景方向。

原文摘要 · Abstract (English)

Physical reasoning over visual inputs demands tight integration of visual perception, domain knowledge, and multi-step symbolic inference. Yet even state-of-the-art Vision Language Models (VLMs) fall far short of human performance on physics benchmarks. While post-training algorithms such as Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO) have demonstrated strong reasoning gains in language models, how reward design shapes VLM physical reasoning behavior remains poorly understood. We present a systematic reward ablation study for GRPO-based VLM training on physical reasoning. We compare four reward signals of increasing semantic richness: format compliance, answer accuracy, a composite rubric reward (answer correctness, physics principle identification, and unit consistency), and a novel internal reward derived from model attention weights over input image regions. We evaluate on PhyX, a 3,000-problem benchmark spanning six physics domains and six reasoning types across multiple-choice and open-ended formats, using IBM Granite Vision 3.3 (2B). Across both formats, GRPO with accuracy-based rewards outperforms SFT on most domains, though gains vary substantially by reward type and domain. Reward design does not uniformly improve performance. Instead, it induces domain-specific reasoning behaviors. Accuracy-based rewards provide the strongest overall gains. Rubric rewards improve structured reasoning quality without consistent accuracy improvements. Attention-based rewards enhance spatial reasoning while degrading performance in symbolic domains. Our internal attention-weight reward requires no spatial annotations and improves spatial relation accuracy from 0.27 to 0.50, suggesting that supervising where the model attends during generation is a promising direction for visually grounded physical reasoning.

物理推理奖励设计视觉语言模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。