arXiv:2601.22532cs.LGcs.AI2026-01

从批量上下文博弈视角解析强化微调的设计选择,揭示关键因素。

Demystifying Design Choices of Reinforcement Fine-tuning: A Batched Contextual Bandit Learning Perspective

  • 构建最小化基线,分离各设计因素影响
  • 实验证明优势计算与滚展开次数是关键变量
  • 适合关注强化微调机制的算法研究者

强化微调领域论文爆发式增长,多聚焦于优化设计选择。尽管性能提升常被宣称,但结论不一致,进展显得模糊。反思这一现象,我们仍缺乏对两个根本问题的严谨回答:1)每个设计选择的作用是什么?2)哪些是关键因素?本文旨在解答。核心挑战在于设计选择相互纠缠,难以准确归因其对学习和泛化的影响。为此,我们构建了一个最小化基线:每轮仅一次采样,以回报作为训练信号且不使用优势技巧,批次大小为32。该基线与批量上下文博弈学习相联系,便于实验分析。围绕此基线,设计实验流程,评估优势计算、采样次数等因子的边际收益。在三个基础模型和两个数据集上的实验不仅揭示了不同设计选择对学习与泛化动态的影响,还识别出需重点投入的关键因素。

原文摘要 · Abstract (English)

The reinforcement fine-tuning area is undergoing an explosion papers largely on optimizing design choices. Though performance gains are often claimed, inconsistent conclusions also arise from time to time, making the progress illusive. Reflecting on this illusion, we still lack principled answers to two fundamental questions: 1) what is the role of each design choice? 2) which ones are critical? This paper aims to shed light on them. The underlying challenge is that design choices are entangled together, making their contribution to learning and generalization difficult to attribute. To address this challenge, we first construct a minimalist baseline for disentangling factors: one rollout per query in each round, the outcome reward serving as the training signal without any advantage trick, and a batch size of thirty-two. This baseline connects to batched contextual bandit learning, which facilitates experimental analysis. Centering around this baseline, we design an experiment pipeline, examining the marginal gains of factors like advantage, number of rollouts, etc. Experiments on three base models and two datasets, not only reveal new understanding on the role of various design choices on learning and generalization dynamics, but also identify critical ones that deserve more effort.

强化微调实验分析设计选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。