arXiv:2607.18110cs.LGcs.CL2026-07被引 1

让大模型当教练,用经验知识提升开放任务学习效果

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

  • 用大模型做教练,将评估反馈转化为可迁移的经验知识
  • 相比单一奖励信号,新方法在多个任务上表现更优且泛化更强
  • 适合需要精细判断的开放性任务,如创意生成与复杂推理

在开放性任务中,强化学习将大模型的评分标准压缩为单一奖励信号,丢失了丰富的文本反馈并混淆了不同质量的回答。我们提出体验式学习(EL),将大模型评判系统转为大模型教练。教练将每个策略响应的评估结果提炼为可迁移的体验知识,指导教师模型,并通过在线上下文蒸馏被策略内部化。相比标量奖励,这种高带宽反馈通道提供密集监督,保留高质量回答间的细微偏好。在两种策略家族中,使用自评或专有模型反馈,EL 在保留和未见的开放性任务上均持续优于基于评分的强化学习。显著的是,EL 能更好外推至训练分布之外,缓解奖励欺骗问题。这些结果确立体验知识作为后训练阶段非可验证任务中更丰富、更通用的学习信号。

原文摘要 · Abstract (English)

Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach. The coach distills its assessment of each on-policy response into transferable experiential knowledge, which conditions a teacher model and is internalized by the policy through on-policy context distillation. Compared with scalar rewards, this higher-bandwidth feedback channel provides dense supervision and preserves fine-grained preferences among high-quality responses. Across two policy families, with feedback from the policy itself or a proprietary model, EL consistently outperforms rubric-based RL on held-out and unseen open-ended tasks. Notably, EL generalizes better beyond the training distribution, and mitigates reward hacking. These findings establish experiential knowledge as a richer and more generalizable learning signal for post-training on non-verifiable tasks.

大模型教练体验学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。