小模型上用GRPO优化反而让网页代理表现变差,说明它只在有提升空间时才有效。
A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism
- 在4B到8B规模的小模型上测试GRPO,发现多数配置无法提升成功率
- 中高学习率使文本任务成功率显著下降,且结果具有统计显著性
- 失败源于学习率过高导致模型崩溃或注意力层退化,仅在小模型中明显
在4B至8B规模的小型语言与视觉-语言模型网页代理中,我们系统评估了基于可验证奖励的强化学习方法,尤其是组相对策略优化(GRPO)的有效性。通过18组对照实验,改变学习率、KL权重、随机种子、初始化和裁剪策略,未发现任何配置能显著提升强监督基线在已掌握任务上的成功率。在文本任务上,中高学习率甚至导致成功率显著下降。该负效果在配对测试、25个评估种子、6个训练种子、不同训练配方、文本与标记集截图观测输入以及8B模型扩展下均持续存在;只有在标记集任务中为轻微恶化。为排除流程问题,我们在采样可达奖励的任务上使用相同框架,成功率达22个百分点提升且置信区间不包含零,表明GRPO仅在存在提升空间时有效。进一步分析揭示:中等学习率引发局部退化(集中于注意力与MLP模块),而高学习率导致全局崩溃,且嵌入层变化虽主导权重移动但无因果作用。在4B模型中,深层有效秩与性能双向相关;在8B模型中两者分离。该耦合关系具有尺度依赖性。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent. We ask whether it adds skill to a small language and vision-language model web agent at the 4B to 8B scale, or whether it mostly reshapes behavior the supervised model already has. Across a control grid of 18 runs that varies learning rate, KL weight, seed, initialization, and clipping, no configuration credibly improves the success rate of a strong supervised baseline on tasks the agent has largely mastered. On the text track, moderate to high learning rates make it credibly worse. The null holds under paired testing, 25 evaluation seeds, 6 training seeds, changes to the recipe, both text and Set-of-Marks screenshot observations, and scaling the backbone to 8B; the credible harm is a text-track finding and is only nominal under Set-of-Marks. To show that the null reflects the setting and not a broken pipeline, we run the identical harness, reward, and recipe on tasks whose reward is reachable by sampling, and there the success rate rises by 22 points with a paired interval that excludes zero. GRPO therefore helps only when there is headroom to climb, meaning the sampled policy already succeeds more often than the greedy one. We then explain the failure. A middle learning rate degrades the agent and a high one collapses it, and the two regimes form a double dissociation: grafting localizes the degrade regime to the attention and MLP blocks, while the collapse regime cannot be traced to any single group, and the embedding change that dominates the weight movement is causally inert. At 4B, effective rank in the late layers tracks capability in both directions; at 8B the two come apart. This coupling is specific to the smaller model, so we report it as scale-dependent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。