arXiv:2606.08625cs.CL2026-06被引 2

用评分标准统一评估大模型,让人类意图可被机器理解

From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape

论文配图:From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape
图 1 · 摘自论文原文
  • 提出'评分标准'框架,将模糊评价转为可执行规则
  • 三层次应用:评价分解、训练反馈、模型自优化
  • 适合关注模型对齐与评估可靠性的研究者

随着大语言模型向开放式自主代理演进,评估与引导其行为的机制也需同步发展。本文提出评分标准作为统一框架,描述其作为对大模型范式演进的动态响应,在评估、强化学习与安全对齐等独立研究中反复出现,具有非偶然性。评分标准是显式的评判准则集合,能将复杂的质量判断转化为结构化、可操作的标准。我们系统梳理现有评分标准设计,分析其构建与优化方式,并考察其在评估与训练中的作用。评分标准体现为三个递进层次:评价层将整体判断拆解为可验证维度;训练层提供细粒度反馈信号,弥补标量奖励的不足;内在层则从模型行为中动态生成,驱动自我改进。我们进一步评估评分标准在生成质量、执行保真度、理论约束与安全威胁下的可靠性,并综述跨领域的基于评分标准的基准测试。通过使评估透明且可分解,评分标准将人类价值期望转化为机器可学习信号,成为连接人类意图与机器行为的持久桥梁。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) advance toward open-ended autonomous agents, the mechanisms used to evaluate and guide their behavior must evolve accordingly. This work introduces the rubric as a unifying framework capturing this evolution, characterizing rubrics as a dynamic response to successive LLM paradigm shifts that recurs across otherwise independent efforts in evaluation, reinforcement learning, and safety alignment. We define rubrics as explicit criteria sets that transform complex quality judgments into structured and actionable standards, and demonstrate that their recurrence across these research threads is not coincidental. We systematically organize existing rubric designs, examine their construction and optimization, and analyze their role across evaluation and training. Rubrics manifest at three progressively deeper levels: at the evaluative level, they decompose holistic judgments into verifiable dimensions; at the training level, they serve as dense feedback signals providing process-level guidance where scalar rewards fall short; at the intrinsic level, they emerge dynamically from model behaviors, driving self-improvement. We further assess rubric reliability across generation quality, execution fidelity, theoretical constraints, and security threats, before surveying rubric-based benchmarks across diverse domains. By rendering assessment transparent and decomposable, rubrics translate human value expectations into machine-learnable signals, serving as the enduring bridge between human intentions and machine behavior.

大模型评估评分标准对齐机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。