arXiv:2605.29496cs.CLcs.CV2026-05

发现视觉语言模型后训练中推理强于感知,提出改进方法提升感知能力。

On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training

论文配图:On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
图 1 · 摘自论文原文
  • 通过合成任务分离感知与推理,诊断出训练中的不对称问题。
  • 动态重加权损失使端到端性能提升18.2点,感知奖励改进6.0点。
  • 适合关注多模态模型平衡优化的研究者与开发者。

后训练显著提升了前沿视觉语言模型的推理能力,但对感知能力的提升相对有限,成为端到端视觉推理的瓶颈。为探究这一差距,我们引入一个受控诊断框架,包含两个解耦感知与推理的合成任务。分析揭示出一致的感知-推理不对称性:后训练对推理的提升远大于感知,且不同训练范式机制不同。在监督微调(SFT)中,这种不对称源于思维链监督中的标记不平衡,感知占据更少标记,导致训练信号较弱;动态重加权损失可缓解此问题,使端到端性能提升高达18.2。在强化学习(RL)中,不对称源于奖励耦合:结果奖励与推理相关性更强,削弱了对感知的信号;引入感知感知奖励可缓解此问题,提升端到端准确率最高达6.0;即使无真实感知奖励,可靠的替代奖励也能提供有效信号,带来3.2点增益。结果全面诊断了不对称优化现象,并提出了具体干预策略以平衡感知与推理。

原文摘要 · Abstract (English)

Post-training has greatly improved reasoning in frontier vision-language models, yet its gains for perception remain comparatively limited, creating a bottleneck for end-to-end visual reasoning. To investigate this gap, we introduce a controlled diagnostic framework with two synthetic tasks that disentangle perception from reasoning. Our analysis reveals a consistent perception-reasoning asymmetry: posttraining improves reasoning more substantially than perception, though the underlying mechanism differs by training paradigm. For supervised fine-tuning (SFT), this asymmetry stems from token imbalance in chain-of-thought supervision, where perception occupies fewer tokens and thus receives a weaker training signal. Dynamically reweighting the loss mitigates this imbalance and boosts end-to-end performance by up to 18.2. For reinforcement learning (RL), the asymmetry instead arises from reward coupling: outcome rewards correlate more strongly with reasoning than with perception, weakening the signal for perception learning. Adding a perception-aware reward alleviates the imbalance and improves end-to-end accuracy by up to 6.0; even without groundtruth perception rewards, a reliable surrogate reward provide useful signal, yielding gains of 3.2 points. Together, our results comprehensively diagnose asymmetric optimization and suggest concrete interventions to balance perception and reasoning.

视觉语言模型后训练感知推理平衡强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。