首次为百亿参数大模型的强化学习提供可验证奖励的非平凡泛化界。
Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards

- 用压缩的PAC-Bayes方法结合门控最大重参数化,应对生成随机性。
- 在4个任务中性能保留84%-97%,模型压缩率提升14,796倍。
- 适合关注大模型泛化能力与高效部署的研究者。
尽管基于可验证奖励的强化学习(RLVR)被广泛用于提升大语言模型(LLMs)的推理能力,但其模型泛化能力仍不清晰。本文首次建立了百亿参数规模下参数高效RLVR微调的非平凡泛化边界。方法上,将PAC-Bayes压缩界适配至该场景,并通过Gumbel-max重参数化处理标记生成的固有随机性。为实现这些边界,提出Progressive RLVR框架,融合了RLVR、在线策略蒸馏、TinyLoRA与模型量化。该框架在保持标准LoRA微调84%-97%性能的同时,使模型压缩率提升14,796倍。在数学求解、编程、通用知识推理和Text-to-SQL四个领域,所得边界优于基础模型准确率9%-51%,且接近微调模型准确率的6%-11%。
原文摘要 · Abstract (English)
While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the resulting models remains poorly understood. In this work, we establish the first non-vacuous generalization bounds for parameter-efficient RLVR fine-tuning at the billion-parameter scale. Our approach adapts PAC-Bayes compression bounds to this setting, and addresses the inherent stochasticity of token generation by applying the Gumbel-max reparameterization trick. To operationalize these bounds, we propose the Progressive RLVR framework, which integrates RLVR with on-policy distillation, TinyLoRA, and model quantization. Progressive RLVR empirically retains 84-97% performance of standard LoRA fine-tuning while producing models that are 14,796x more compressible. We show that this framework yields non-vacuous generalization bounds in four domains: mathematical problem-solving, programming, general-knowledge reasoning, and Text-to-SQL. Our bounds exceed the accuracy of the base model by 9-51% and lie within 6-11% of the accuracy of the fine-tuned models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。