arXiv:2601.14599cs.LGcs.AI2026-01

用博弈论视角重新审视大模型强化微调的优化策略

Rethinking Reinforcement fine-tuning of LLMs: A Multi-armed Bandit Learning Perspective

  • 构建极简实验框架,将微调视为大规模离散动作的多臂赌博机问题
  • 在3个大模型、2个推理数据集上验证,发现奖励直接作为信号更有效
  • 揭示各优化设计的真实作用,适合想深入理解微调机制的研究者

大量启发式方法被提出用于优化大语言模型的强化微调,但时常出现相互矛盾的结论,使该领域难以把握。针对这一困境,本文聚焦两个根本性问题:1)每个优化选择的实际作用是什么?2)哪些是关键瓶颈?为此,我们提出自下而上的实验流程。底层采用极简配置:单一训练数据、每轮仅一次回溯采样,且奖励直接作为学习信号,不使用优势函数。该配置对应于具有极大离散动作空间的多臂赌博机学习问题,为实验结果提供理论支持。实验流程逐层扩展该极简配置,系统检验每个设计选择的影响。在三个大语言模型和两个推理数据集上的实验不仅揭示了对设计选择的新认知,还为该领域提供了关键洞见。

原文摘要 · Abstract (English)

A large number of heuristics have been proposed to optimize the reinforcement fine-tuning of LLMs. However, inconsistent claims are made from time to time, making this area elusive. Reflecting on this situation, two fundamental questions still lack a clear understanding: 1) what is the role of each optimizing choice? 2) which ones are the bottlenecks? This paper aims to shed light on them, and it faces the challenge of several entangled confounding factors in the fine-tuning process. To tackle this challenge, we propose a bottom-up experiment pipeline. The bottom layer is composed of a minimalist configuration: one training data, one rollout per round and the reward directly serve as the learning signal without advantage function design. This minimalist configuration connects to multi-armed bandit learning with extremely large discrete action space, which offers theories to corroborate the experiment findings. The up procedure of the experiment pipeline expanding the minimalist configuration layer by layer, examining the role of each design choice. Experimental results on three LLMs and two reasoning datasets not only reveal new understanding of the design choice but also yield essential insights to shape the area.

强化微调大模型实验分析多臂赌博机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。