揭示了在可验证奖励下组合推理能否学习的理论条件。
When Is Compositional Reasoning Learnable from Verifiable Rewards?
- 提出任务优势比衡量组合问题与模型的适配性。
- 中间步骤正确能带来明显优势时,组合推理可高效学习。
- 适合研究强化学习与大模型推理机制的学者参考。
大型语言模型通过可验证奖励的强化学习(RLVR)实现组合推理的涌现,是近期实证成功的关键驱动力。然而,在仅使用结果级反馈的情况下,哪些组合问题可学习仍不明确。本文从理论上研究了自回归模型在RLVR训练下的组合问题可学习性。我们提出一个称为任务优势比的量,它是组合问题与基础模型的联合属性,用于刻画哪些任务和组合结构可从结果级反馈中学习。正面结果表明,当正确中间步骤能带来明显优势时,此类组合问题可通过RLVR高效学习;我们还分析了该优势在不同问题中的自然产生机制。负面结果则指出,若缺乏结构性优势,RLVR可能收敛至次优组合。我们证明,在某些情况下,基础模型的质量决定了该优势是否存在,以及是否会导致次优解。本研究旨在为RLVR何时成功、何时失败提供原则性理论理解。
原文摘要 · Abstract (English)
The emergence of compositional reasoning in large language models through reinforcement learning with verifiable rewards (RLVR) has been a key driver of recent empirical successes. Despite this progress, it remains unclear which compositional problems are learnable in this setting using outcome-level feedback alone. In this work, we theoretically study the learnability of compositional problems in autoregressive models under RLVR training. We identify a quantity that we call the task-advantage ratio, a joint property of the compositional problem and the base model, that characterizes which tasks and compositions are learnable from outcome-level feedback. On the positive side, using this characterization, we show that compositional problems where correct intermediate steps provide a clear advantage are efficiently learnable with RLVR. We also analyze how such an advantage naturally arises in different problems. On the negative side, when the structural advantage is not present, RLVR may converge to suboptimal compositions. We prove that, in some cases, the quality of the base model determines if such an advantage exists and whether RLVR will converge to a suboptimal solution. We hope our analysis can provide a principled theoretical understanding of when and why RLVR succeeds and when it does not.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。