强化后训练提升推理能力,但跨领域泛化效果不稳定。
Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?
- 对比多个开源RPT模型与基础模型在多领域表现
- 跨领域测试中,推理增益常消失或不一致
- 适合关注LLM泛化能力的研究者阅读
强化后训练(RPT)近期被证明可提升大语言模型(LLMs)的推理能力。然而,现有研究多在与微调数据相同领域评估,其跨领域泛化能力尚不明确。本文通过两项研究考察基于可验证奖励的强化学习(RLVR)的泛化性:(1) 观察性研究:在多个领域(含已见和未见领域)比较多种开源RPT模型与其基线模型;(2) 干预性研究:仅在一个领域上对LLM进行RPT微调,再评估其在多领域上的表现。两项研究均表明,尽管RPT在与微调数据相似的任务上带来显著提升,但在推理模式不同的新领域上,性能增益往往无法保持甚至消失。
原文摘要 · Abstract (English)
Reinforcement post training (RPT) has recently shown promise in improving the reasoning abilities of large language models (LLMs). However, it remains unclear how well these improvements generalize to new domains, as prior work evaluates RPT models on data from the same domains used for post-training. To understand the generalizability of RPT, we conduct two studies with specific focus on Reinforcement Learning with Verifiable Rewards (RLVR). (1) Observational: we compare a wide range of open-weight RPT models against their corresponding base models across multiple domains, including both seen and unseen domains in their fine-tuning data. (2) Interventional: we fine-tune LLMs with RPT on single domains and evaluate their performance across multiple domains. Both studies converge on the same conclusion that, although RPT brings substantial gains on tasks similar to the fine-tuning data, the gains generalize inconsistently and can vanish on domains with different reasoning patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。