提升大模型推理多样性,避免答案趋同。
SetPO: Set-Level Policy Optimization for Diversity-Preserving LLM Reasoning
- 用核函数衡量轨迹相似性,从整体层面优化多样性。
- 在多个基准上,通过率高于基线模型,且效果随规模提升。
- 适合需要多角度解题的数学推理任务研究者。
基于可验证奖励的强化学习在提升大语言模型(LLMs)推理能力方面表现优异,尤其在数学任务中。但此类方法常导致结果多样性下降,模型概率集中于少数解法。受收益递减原理启发,我们提出一种基于采样轨迹的集合级多样性目标,采用核函数计算轨迹间相似性。通过推导每条轨迹的留一法边际贡献,并将其作为即插即用的优势调节项融入策略优化。进一步在分布扰动框架下分析单条轨迹对模型多样性的贡献,理论证明了稀有轨迹对全局多样性的边际贡献具有单调性。在多种模型规模和基准测试上的实验表明,所提算法在 Pass@1 与 Pass@K 指标上均持续优于强基线。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards has shown notable effectiveness in enhancing large language models (LLMs) reasoning performance, especially in mathematics tasks. However, such improvements often come with reduced outcome diversity, where the model concentrates probability mass on a narrow set of solutions. Motivated by diminishing-returns principles, we introduce a set level diversity objective defined over sampled trajectories using kernelized similarity. Our approach derives a leave-one-out marginal contribution for each sampled trajectory and integrates this objective as a plug-in advantage shaping term for policy optimization. We further investigate the contribution of a single trajectory to language model diversity within a distribution perturbation framework. This analysis theoretically confirms a monotonicity property, proving that rarer trajectories yield consistently higher marginal contributions to the global diversity. Extensive experiments across a range of model scales demonstrate the effectiveness of our proposed algorithm, consistently outperforming strong baselines in both Pass@1 and Pass@K across various benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。