Autorubric统一评估框架,解决大模型非可验证任务评估中的偏见与不一致问题。
Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks
- 设计可复用、可审计的评估框架,显式管理评分标准与判断逻辑
- 在化学、研究系统等任务中发现评分准则特异性失效和模型家族差异
- 支持评分解释与强化学习奖励建模,适合评估与优化研究者
基于评分量表的大模型评估已成为非可验证任务(无法通过程序化判定)评估与优化的关键工具。然而现有评估机制仍面临位置偏倚、随机不一致、标准混淆、不确定下强制判断及模型依赖校准等问题。评分量表虽结构化评估流程,却引入了关于标准设计、量级类型、权重分配、聚合方式、弃权策略、校准与可靠性度量等关键决策。尽管相关研究分散于大模型评估、教育测量与心理计量领域,方法仍零散且实现不完整,导致研究者重复造轮子。我们提出Autorubric——一个开源框架,使评分与评估选择显式、可复用、可审计。通过统一API,支持混合标准类型的原子评估,集成可配置的偏倚缓解、校准、集成、弃权与心理测量诊断。在大学化学评分、深度研究系统及新提出的混合标准对话机器人基准CHARM-100上的评估揭示了准则特异性失败、系统性评估者族差异及配置效应,表明不存在普适的缓解方案。进一步证明,逐准则得分与解释可支撑智能体技能修正与强化学习奖励建模。Autorubric提供从测量到优化的共享基础设施,为方法比较、评估复现与跨研究证据积累奠定共同基础。
原文摘要 · Abstract (English)
Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduced to exact programmatic checks. Yet the underlying judges remain vulnerable to position bias, stochastic inconsistency, criterion conflation, forced judgments under uncertainty, and model-dependent calibration. Rubrics structure evaluation but do not eliminate these failures; they add consequential choices about criterion design, scale types, weighting, aggregation, abstention, calibration, and reliability measurement. Despite extensive research across LLM evaluation, educational measurement, and psychometrics, the relevant methods remain scattered across papers and partial implementations. Researchers therefore pay a reinvention tax, repeatedly rebuilding evaluation machinery instead of accumulating knowledge on a common substrate. We introduce Autorubric, an open-source framework that makes rubric and judge choices explicit, reusable, and auditable. Through a unified API, it supports atomic evaluation of mixed criterion types alongside configurable bias mitigations, calibration, ensembling, abstention, and psychometric diagnostics. Evaluations spanning college chemistry grading, deep-research systems, and CHARM-100---a new mixed-criterion chatbot benchmark---reveal criterion-specific failures, systematic judge-family differences, and configuration effects that do not support a universal mitigation stack. We further demonstrate how per-criterion scores and explanations can support agent skill revision and reward modeling for reinforcement learning. By providing shared infrastructure from measurement through optimization, Autorubric gives the community a common basis for comparing methods, reproducing evaluation choices, and accumulating evidence across studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。