arXiv:2604.11522cs.CL2026-04ACL

用信息增益修正奖励,让文字生成更丰富不重复。

Triviality Corrected Endogenous Reward

论文配图:Triviality Corrected Endogenous Reward
图 1 · 摘自论文原文
  • 用专家与通用模型的差异作为奖励信号
  • 在多个写作任务上提升生成质量且无需外部标注
  • 适用于文字生成和数学推理,通用性强

开放性文本生成中的强化学习受限于无法验证的奖励,通常依赖需标注数据或闭源模型的评判系统。受基于置信度的无监督强化学习启发,我们探索将此原则用于开放性写作任务。发现直接使用置信度奖励会导致平凡性偏差:策略坍缩至高概率输出,降低多样性与内容意义。为此提出TCER(平凡性修正内生奖励),通过奖励专家策略与通用参考策略之间的相对信息增益,并引入概率依赖校正机制来缓解该偏差。在多个写作基准和模型架构上,TCER均实现一致改进,且无需外部监督。此外,该方法在数学推理任务中也表现出良好迁移能力,验证了其跨生成任务的通用性。

原文摘要 · Abstract (English)

Reinforcement learning for open-ended text generation is constrained by the lack of verifiable rewards, necessitating reliance on judge models that require either annotated data or powerful closed-source models. Inspired by recent work on unsupervised reinforcement learning for mathematical reasoning using confidence-based endogenous rewards, we investigate whether this principle can be adapted to open-ended writing tasks. We find that directly applying confidence rewards leads to Triviality Bias: the policy collapses toward high-probability outputs, reducing diversity and meaningful content. We propose TCER (Triviality Corrected Endogenous Reward), which addresses this bias by rewarding the relative information gain between a specialist policy and a generalist reference policy, modulated by a probability-dependent correction mechanism. Across multiple writing benchmarks and model architectures, TCER achieves consistent improvements without external supervision. Furthermore, TCER also transfers effectively to mathematical reasoning, validating the generality of our approach across different generation tasks.

强化学习文本生成内生奖励无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。