arXiv:2608.27505cs.CLcs.AI2026-08中稿 · EMNLP综述

用结构化评分标准提升大模型对齐效果,让奖励更可解释、更可靠。

A Survey on Rubric-Guided Reinforcement Learning for Language Models

  • 基于贝叶斯框架,将评分标准视为先验分布,实例化为条件评分
  • 提出五类评分引导强化学习方法,覆盖从静态到自演化全过程
  • 适合研究大模型对齐、奖励设计及可解释性方向的学者阅读

基于人类反馈的强化学习(RLHF)已成为对齐大型语言模型与人类偏好主流范式。然而传统RLHF依赖标量奖励信号,缺乏可解释性且难以捕捉响应质量的多维特性。评分引导强化学习通过引入结构化、可解释的评估标准(即评分标准)作为奖励设计、反馈生成和策略优化的核心,克服了上述局限。本文提出一个贝叶斯框架,将‘宪法’定义为评价标准上的先验分布 $P(R)$,而评分标准则是给定输入 $x$ 的条件实例化 $R_x ilde{P}(R|x)$。在此统一视角下,我们沿先验-后验轴构建了评分引导强化学习的分类体系,涵盖宪法AI、实例特定评分标准、过程级监督、自演化评分标准及其智能体与多模态扩展。此外,由于评分标准是自然语言产物,我们还分析了粒度权衡、语义漂移与语言奖励劫持如何影响对齐可靠性,指出了未来研究的关键开放问题。

原文摘要 · Abstract (English)

Reinforcement learning from human feedback (RLHF) has become the dominant paradigm for aligning large language models (LLMs) with human preferences. However, traditional RLHF relies on scalar reward signals that lack interpretability and fail to capture the multifaceted nature of response quality. Rubric-guided reinforcement learning addresses these limitations by introducing structured, interpretable evaluation criteria, or rubrics, as the backbone of reward design, feedback generation, and policy optimization. In this survey, we introduce a Bayesian framework that defines constitutions as prior distributions $P(R)$ over evaluation criteria and rubrics as conditional instantiations $R_x \sim P(R|x)$. Under this unified view, we present a taxonomy of rubric-guided RL along the prior-posterior axis, covering constitutional AI, instance-specific rubrics, process-level supervision, self-evolving rubrics, and their agentic and multimodal extensions. Furthermore, as rubrics are natural-language artifacts, we present a linguistic analysis of how granularity trade-offs, semantic drift, and linguistic reward hacking impact alignment reliability, identifying key open problems for future research.

大模型对齐强化学习评分标准可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。