用结果层面强化学习提升模型对新组合的泛化能力
Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization

- 以最终输出为反馈,通过强化学习优化模型
- 在多个基准上优于监督微调,尤其提升复杂组合泛化
- 适合需要处理未知组合场景的研究者
组合泛化指正确理解已知基本单元的新组合,仍是重大挑战。现有方法多依赖监督微调,导致模型仅模仿目标输出,无法捕捉全局组合结构。本文探索是否可通过结果层面的强化学习改进组合泛化。采用组相对策略优化,基于最终输出反馈进行模型训练,设计了简单二元奖励和含组合反馈的复合奖励。在多个组合基准上的实验表明,强化学习显著优于监督微调。进一步分析显示,监督模型易过拟合高频训练组合,而强化学习通过重塑输出分布,尤其提升了复杂组合类型的泛化能力。
原文摘要 · Abstract (English)
Compositional generalization refers to correctly interpret novel combinations of known primitives, which remains a major challenge. Existing approaches often rely on supervised fine-tuning, which encourages models to imitate target outputs. This token-level training paradigm fails to capture the global compositional structure required for generalizing to unseen combinations. In this work, we investigate whether compositional generalization can instead be improved through outcome-level reinforcement learning. We adopt Group Relative Policy Optimization to optimize models based on feedback on their final outputs. Within this framework, we explore both a simple binary outcome reward and a composite reward that provides additional composition feedback. Experiments on multiple compositional benchmarks show that reinforcement learning improves compositional generalization compared to supervised fine-tuning. Further analysis reveals that supervised models tend to overfit frequent training compositions, whereas reinforcement learning improves compositional generalization by reshaping the output distribution, particularly for more complex composition types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。