arXiv:2508.05428cs.LG2025-08被引 3

让大模型生成更协调,通过因果分析优化多候选回复的协同关系。

Group Causal Policy Optimization for Post-Training Large Language Models

  • 引入因果模型揭示回复间的隐含依赖关系
  • 在多个推理基准上超越GRPO等现有方法
  • 适合需要高质量多回复协同的后训练场景

大语言模型在多样化任务中表现优异,但特定领域仍需针对性后训练。现有方法中,组相对策略优化(GRPO)因高效而突出,其利用组间相对奖励避免昂贵的价值函数学习,但将候选回复视为独立,忽略了互补与矛盾等语义交互。为此,我们首先构建结构因果模型(SCM),揭示在最终整合输出条件下候选回复间因形成碰撞器结构而产生的隐藏依赖。因果分析得出两项洞见:(1) 将回复投影至因果引导子空间可提升预测质量;(2) 此投影生成的基线优于仅基于查询的条件化。基于此,我们提出组因果策略优化(GCPO),通过两个核心组件整合因果结构:因果引导的奖励调整,以及对齐于因果投影参考分布的新式KL正则项。全面实验表明,GCPO在多个推理基准上持续优于现有方法,包括GRPO。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have broadened their applicability across diverse tasks, yet specialized domains still require targeted post training. Among existing methods, Group Relative Policy Optimization (GRPO) stands out for its efficiency, leveraging groupwise relative rewards while avoiding costly value function learning. However, GRPO treats candidate responses as independent, overlooking semantic interactions such as complementarity and contradiction. To address this challenge, we first introduce a Structural Causal Model (SCM) that reveals hidden dependencies among candidate responses induced by conditioning on a final integrated output forming a collider structure. Then, our causal analysis leads to two insights: (1) projecting responses onto a causally informed subspace improves prediction quality, and (2) this projection yields a better baseline than query only conditioning. Building on these insights, we propose Group Causal Policy Optimization (GCPO), which integrates causal structure into optimization through two key components: a causally informed reward adjustment and a novel KL regularization term that aligns the policy with a causally projected reference distribution. Comprehensive experimental evaluations demonstrate that GCPO consistently surpasses existing methods, including GRPO across multiple reasoning benchmarks.

因果推理大模型训练强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。