用评分标准设计奖励机制,平衡长文本生成的准确性与丰富性。
From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning
- 以关键点评分表定义回答应包含内容,直接作为奖励信号。
- 严格惩罚幻觉提升准确率,但会压缩信息量;纯评分表则增强丰富性但降低可信度。
- 融合三类奖励可最优平衡,更适应分布外任务。
仅惩罚无依据陈述虽能提升生成内容的准确性,却可能使模型回答更简略。本文研究长文本生成中“拒绝-丰富”权衡问题。不采用长度、论点数或相关性等全局丰富性代理指标,而是为每个问题设计关键点评分表,明确回答所需及可选信息。该评分表既用于评估,也作为奖励信号。实验对比了仅接地、代理型、仅评分表、以及混合奖励四种方式,发现存在稳定权衡:严格接地奖励提升支持度但抑制覆盖范围,而无约束评分表奖励提升覆盖度但削弱接地性。软性结合接地性、评分表覆盖率和相关性,在分布内任务中显著提升支持度,并在分布外检查表任务上优于单一奖励策略。
原文摘要 · Abstract (English)
Rewards that penalize unsupported claims can improve grounding in long-form generation, but they can also teach models to answer less. We study this refusal-to-richness trade-off in long-form hallucination RL. Instead of using global richness proxies such as length, claim count, detail, or pairwise relevance, we represent each question with a key-point rubric that specifies the required and optional information a useful answer should cover. These rubrics define coverage directly and are used both for evaluation and as reward signals. Across grounding-only, proxy-based, rubric-only, and combined rewards, we find a stable trade-off: strict grounding rewards improve support but suppress coverage, while unconstrained rubric rewards improve coverage but weaken grounding. A soft combination of grounding, rubric coverage, and relevance gives the best balance in our experiments, improving in-distribution support while transferring better to out-of-distribution checklist tasks than either grounding-only or rubric-only rewards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。