arXiv:2604.01375cs.AI2026-04被引 2

为评分标准的失效模式建立系统分类,助力自动诊断评分问题。

RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics

  • 构建八类失效模式的分类体系,涵盖可靠性、内容有效性等三方面。
  • 人工标注一致率达87%,卡帕系数0.64,验证分类稳定性。
  • 开发自动化指标,与人工标注匹配度高达0.925 F1,支持大规模诊断。

基于评分标准的评估广泛应用于大模型在开放性、难以验证任务中的基准测试与训练流程。尽管已有研究通过下游信号(如强化学习结果)证明评分标准的有效性,但尚无系统方法仅从这些聚合或下游信号中诊断评分标准本身的失效原因。为此,我们提出RIFT:RubrIc Failure mode Taxonomy,一个用于系统化刻画评分标准构建与设计中失效模式的分类体系。RIFT包含八种失效模式,分为三大类别:可靠性失效、内容效度失效与后果效度失效。该分类基于扎根理论,通过对来自五个不同数据源(涵盖通用指令遵循、代码生成、创意写作及专家级深度研究)的评分标准进行迭代标注,直至不再识别出新失效模式为止。我们通过独立人工标注者的一致性评估分类体系,整体达到87%的两两一致性与0.64的平均科恩卡帕系数。最后,为支持可扩展诊断,我们提出自动化评分质量度量,并证明其与人工失效模式标注高度一致,最高达0.925 F1。

原文摘要 · Abstract (English)

Rubric-based evaluation is widely used in LLM benchmarks and training pipelines for open-ended, less verifiable tasks. While prior work has demonstrated the effectiveness of rubrics using downstream signals such as reinforcement learning outcomes, there remains no principled way to diagnose how a rubric itself fails from such aggregated or downstream signals alone. To address this gap, we introduce RIFT: RubrIc Failure mode Taxonomy, a taxonomy for systematically characterizing failure modes in rubric composition and design. RIFT consists of eight failure modes organized into three high-level categories: Reliability Failures, Content Validity Failures, and Consequential Validity Failures. RIFT is developed using grounded theory by iteratively annotating rubrics drawn from five diverse data sources spanning general instruction following, code generation, creative writing, and expert-level deep research, until no new failure modes are identified. We evaluate the consistency of the taxonomy by measuring agreement among independent human annotators, observing fair agreement overall (87% pairwise agreement and 0.64 average Cohen's kappa). Finally, to support scalable diagnosis, we propose automated rubric quality metrics and show that they align with human failure-mode annotations, achieving up to 0.925 F1.

评分标准失效分析自动化诊断LLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。