在视觉语言数学推理中,通过注入提示、教师监督和价值批评,显著提升稀疏奖励下的强化学习效果。
Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning

- 引入文本提示、教师蒸馏和价值批评作为先验知识,增强策略学习。
- 使用提示引导探索使准确率提升14.4点,且跨域迁移性能显著改善。
- 发现传统评估集存在误导性,最难样本更真实反映模型泛化能力。
在包含20,830个视觉数学问题的池中,Qwen2-VL-2B仅3.6%的回溯路径正确,85-97%的GRPO回溯组完全错误,贡献零梯度。在相同条件下对比十一种方法,分别注入文本提示、分布蒸馏(7B教师)和价值预训练批评者(采用MSE或HL-Gauss分类损失)。当先验能有效传递时,六种方法与其余五种(无先验或先验被教师限制、门控抑制、参数错误的价值批评者)形成明显分离,无论在整体域指标还是跨域迁移(DynaMath)上均表现优异。关键发现在于评估:一个长期用于通用分布检验的子集与真实跨域迁移呈强负相关(斯皮尔曼秩相关rho = -0.74,n=11,置换p=0.011),而最难题目子集则高度正相关(rho = +0.89,p<0.001)。归因于近随机的多项选择子集会奖励“未改变”的模型,导致最优跨域方法表现平庸,最差反而看似领先。结果显示,提示引导探索是提示增益的核心机制,将批评者的MSE损失替换为HL-Gauss交叉熵可带来+14.4分的域内提升。所有准确率均为盲评,采用配对精确检验。
原文摘要 · Abstract (English)
Reinforcement learning for vision-language math reasoning starves under sparse reward: on a pool of 20,830 visual-math problems where Qwen2-VL-2B answers 3.6% of rollouts correctly, 85-97% of GRPO rollout groups are entirely wrong and contribute zero gradient. We train eleven methods under identical conditions in this regime, each injecting a different prior: text (reference-solution hints), distribution (on-policy distillation from a 7B teacher), and value (a value-pretrained critic with an MSE or HL-Gauss categorical loss). A prior helps exactly when it is delivered: the six arms whose prior effectively reaches the policy separate with no overlap from the remaining five -- the no-prior baseline and four arms whose prior is teacher-capped, gated away, or lost to a mis-parameterized critic -- both on the pooled in-domain metric and on cross-domain transfer (DynaMath). The central finding, however, concerns evaluation: one slice of the in-domain pool -- long used as this project's general-distribution check -- anti-correlates with genuine cross-domain transfer (Spearman rho = -0.74, n = 11 arms, permutation p = 0.011), while the hardest in-domain slice predicts it closely (rho = +0.89, p < 0.001). We attribute the inversion to a near-chance multiple-choice subset that rewards models for not having changed; read through it, the best cross-domain method looked mediocre and the worst looked like the champion. Among the methods, hint-guided exploration -- not UFT's auxiliary loss -- drives hint gains, and replacing the critic's MSE loss with HL-Gauss cross-entropy is worth +14.4 points in-domain. All accuracies are blind-judged, with paired exact tests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。