AGPO让大模型在推理时既更准又不缩窄思维,适合工业级精准推理场景。
AGPO: Asymmetric Group Policy Optimization for Verifiable Reasoning and Search Ads Relevance at JD

- 采用负向主导策略抑制错误路径,保留原始模型的探索能力。
- 通过组内方差动态调整正向奖励,聚焦罕见正确路径。
- 在数学基准和京东广告推荐中均实现精度与泛化性能双重提升。
基于可验证奖励的强化学习(RLVR)在提升大语言模型推理能力方面表现优异,但现有方法虽提高了采样效率,却未激发根本性新推理模式,反而使训练后模型的推理边界较基线模型收缩,尤其在大规模采样下基线模型覆盖度更高。本文提出非对称分组策略优化(AGPO),以对抗这一边界收缩现象。AGPO采用负向主导的强化策略,抑制错误推理路径,维持基线模型的探索能力;在正向激励上引入组优势机制,依据组内方差动态放大更新,使模型聚焦于稀有正确路径,同时抑制平凡路径的更新。在五个数学基准上的实验表明,AGPO在保持顶尖准确率的同时,持续提升 pass@$k$ 性能。在京东搜索广告相关性优化的工业级应用中,AGPO显著提升了数据标注质量,推动下游学生模型性能大幅提升。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR) has demonstrated notable success in enhancing the reasoning performance of large language models (LLMs). However, recent studies reveal that while current RLVR methods improve sampling efficiency towards correct paths, they do not elicit fundamentally new reasoning patterns. Instead, the reasoning capability boundary of trained models often narrows compared to their base models, with base models achieving higher coverage at large sample sizes. In this work, we propose Asymmetric Group Policy Optimization (AGPO) to counteract this boundary shrinkage. AGPO adopts a negative-dominant reinforcement strategy to suppress incorrect reasoning paths, maintaining the base model's exploration capacity. For positive reinforcement, AGPO adopts a group advantage mechanism, which scales positive updates based on intra-group variance, allowing the model to focus on rare correct paths while suppressing updates from trivial paths. Our experiments on five mathematical benchmarks demonstrate that AGPO achieves state-of-the-art accuracy while consistently improving pass@$k$ performance at scale. In a large-scale industrial application for search ads relevance optimization, AGPO effectively enhances the quality of the data annotation, leading to substantial performance gains in downstream student models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。