arXiv:2603.01562cs.AI2026-03ACL被引 16

构建首个模型生成评分标准与人工标准对齐的评测基准

RubricBench: Aligning Model-Generated Rubrics with Human Standards

  • 设计1147组对比样本,聚焦复杂输入与表面偏见难题
  • 模型生成评分标准与人工标准存在显著差距,表现明显落后
  • 适合研究大模型评估、对齐与自动评分系统的学者使用

随着大语言模型从简单补全演进到复杂生成任务,奖励模型正逐步转向基于评分标准的评估以减少表层偏差。然而,社区缺乏统一的评测基准来评估这一范式,现有基准既缺乏足够的区分度,也缺少真实评分标准标注。为此,我们提出RubricBench,一个包含1,147组成对比较的精选基准,专门用于评估基于评分标准的评估可靠性。其构建采用多维度筛选流程,聚焦于具有细微输入复杂性和误导性表面偏见的难例,并为每项任务提供由专家标注的原子级评分标准,严格依据指令生成。全面实验表明,模型生成的评分标准与人工标注之间存在显著能力差距,即使最先进的模型在自主制定有效评估标准方面仍严重滞后于人类指导的表现。

原文摘要 · Abstract (English)

As Large Language Model (LLM) alignment evolves from simple completions to complex, highly sophisticated generation, Reward Models are increasingly shifting toward rubric-guided evaluation to mitigate surface-level biases. However, the community lacks a unified benchmark to assess this evaluation paradigm, as existing benchmarks lack both the discriminative complexity and the ground-truth rubric annotations required for rigorous analysis. To bridge this gap, we introduce RubricBench, a curated benchmark with 1,147 pairwise comparisons specifically designed to assess the reliability of rubric-based evaluation. Our construction employs a multi-dimensional filtration pipeline to target hard samples featuring nuanced input complexity and misleading surface bias, augmenting each with expert-annotated, atomic rubrics derived strictly from instructions. Comprehensive experiments reveal a substantial capability gap between human-annotated and model-generated rubrics, indicating that even state-of-the-art models struggle to autonomously specify valid evaluation criteria, lagging considerably behind human-guided performance.

大模型评估评分标准对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。