arXiv:2604.02795cs.CLcs.AI2026-04被引 5

让大模型更懂指令细节,用细粒度打分提升对齐效果

Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks

  • 将粗粒度评分分解为逐词责任判定,实现精准反馈
  • 在多个模型上提升指令与评分层面准确率,优于现有方法
  • 适合需要精细控制输出质量的对话系统研发者

基于评分的强化学习已成为对齐大语言模型与复杂开放域指令任务的有前景方法。然而,现有方法主要依赖响应级奖励,导致奖励稀疏和模糊问题严重。为此,我们提出一种名为Rubrics to Tokens(RTT)的新框架,连接粗粒度响应评分与细粒度词级信用分配。RTT引入词级相关性判别器,预测响应中哪些词对特定约束负责,并通过整合响应级与词级优势的RTT-GRPO优化策略训练模型。此外,在从一维结果奖励转向三维词级评分空间时,我们提出一种新的组归一化方法——样本内词组归一化(Intra-sample Token Group Normalization),以适应这一转变。大量实验与基准测试表明,RTT在不同模型上均一致优于其他基线,在指令级与评分级准确率上表现更优。

原文摘要 · Abstract (English)

Rubric-based Reinforcement Learning (RL) has emerged as a promising approach for aligning Large Language Models (LLMs) with complex, open-domain instruction following tasks. However, existing methods predominantly rely on response-level rewards, introducing severe reward sparsity and reward ambiguity problems. To address these issues, we propose Rubrics to Tokens (RTT), a novel rubric-based RL framework that bridges coarse response-level scores and fine-grained token-level credit assignment. RTT introduces a Token-Level Relevance Discriminator to predict which tokens in the response are responsible for a specific constraint, and optimizes the policy model via RTT-GRPO, which integrates response-level and token-level advantages within a unified framework. Furthermore, when transitioning from one-dimensional, outcome-level reward to three-dimensional reward space in the token-level rubric-based RL, we propose a novel group normalization method, called Intra-sample Token Group Normalization, to accommodate this shift. Extensive experiments and benchmarks demonstrate that RTT consistently outperforms other baselines in both instruction- and rubric-level accuracy across different models.

大模型对齐强化学习评分机制词级反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。