用可复用的评分规则提升开放任务的生成质量
Prompt-Level Reward Specifications for Open-Ended Post-Training

- 将评分标准与计算分离,基于提示构建可复用的评分规则
- 无需人工标注即可实现高质量响应,在多个任务上表现更优
- 适合需要精确约束与整体质量兼顾的开放生成场景
开放任务后训练依赖显式提示相关的成功条件奖励,而非仅依赖事后标量评分。在指令遵循、写作和决策支持任务中,响应质量取决于局部需求、整体偏好和明确约束,但现有奖励方法常使这些标准隐含或仅覆盖有限可验证情形。本文提出一种提示级奖励规范框架,将奖励规范与计算分离。仅需提示,该框架即可离线构建可复用的任务自适应评分标准和可执行硬约束检查器,使奖励标准在训练前显式且跨回滚复用。评分时,以产物为中心的评分标准与代码评分结合独立全局评分,生成归一化的混合奖励,涵盖要求满足度、整体质量和确定性约束。该框架无需人类偏好标注、参考答案或单独训练的奖励模型。实验表明,该奖励可提升离线奖励模型风格排序,并支持多任务上的在线强化学习。消融实验进一步证明评分标准、全局评分和可执行验证提供互补监督。
原文摘要 · Abstract (English)
Open-ended post-training benefits from rewards that make prompt-specific success conditions explicit, rather than relying only on post-hoc scalar scores. In instruction following, writing, and decision-support tasks, response quality depends on local requirements, holistic preferences, and explicit constraints, but existing reward methods often leave these criteria implicit or cover only narrowly verifiable cases. We propose a prompt-level reward specification framework that separates reward specification from reward computation. Given only prompts, our framework constructs reusable task-adaptive rubrics and executable hard-constraint checkers offline, making reward criteria explicit before training and reusable across rollouts. At scoring time, artifact-anchored rubric and code scores are combined with an independent global score for residual holistic quality, yielding a normalized hybrid reward over requirement satisfaction, holistic quality, and deterministic constraints. The framework requires no human preference annotations, reference answers, or a separately trained reward model. Experiments show that the resulting reward improves offline RM-style response ranking and supports online reinforcement learning across multiple open-ended benchmarks. Ablations further show that rubrics, global scoring, and executable verification provide complementary supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。