用自适应评分体系提升大模型开放问答评估的准确性与效率
CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation
- 根据任务类型动态构建评分标准库,结合贝叶斯可测量性过滤
- 在JudgmentBench上将人工黄金标准一致性从0.604提升至0.743
- 仅需49项评分标准即可达到目标相关性,适合高成本评估场景
开放式大模型输出的可靠评估需要细粒度评分标准,但专家制定成本高且难扩展。现有自动化流程依赖严格裁判一致性和二元方差过滤,无法区分可测量与有信息量的标准。我们提出CalibratedRubric,一种任务自适应框架,融合类型特定评分、贝叶斯可测量性过滤和基于项目反应理论(IRT)的评分库构建。该方法通过贝塔-伯努利协议后验估计每项评分标准的可测量性,并采用子模信息覆盖率目标,在观测能力范围内构建紧凑评分库。在金融、医疗、通用和法律基准测试中,可测量性过滤使JudgmentBench上的人工黄金标准一致性从κ=0.604提升至0.743。IRT贪婪选择在所有六个响应块中优于随机选择,且仅需49项而非131项评分标准即可在FinResearchBench决策支持任务中达到目标相关性。任务标签扰动进一步降低系统偏差,验证了任务自适应评分的实际意义。结果表明,CalibratedRubric是一种高效、具备不确定性感知能力的开放式大模型评估方法,其校准效果依赖充分的裁判冗余。
原文摘要 · Abstract (English)
Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge unanimity and binary variance filters, which cannot distinguish measurable rubrics from informative ones. We introduce CalibratedRubric, a task-adaptive framework that combines type-specific scoring, Bayesian rubric-measurability filtering, and item response theory (IRT)-based bank assembly. CalibratedRubric estimates each rubric's measurability with a Beta--Bernoulli agreement posterior and uses a submodular information-coverage objective to construct compact rubric banks over the observed capability range. Across financial, healthcare, general, and legal benchmarks, measurability filtering improves human-gold agreement on JudgmentBench from $κ=0.604$ to $0.743$. IRT-based greedy selection improves cross-fitted rank fidelity over random selection across all six evaluated response blocks and requires only 49 rather than 131 rubrics to reach the target correlation on FinResearchBench decision-support tasks. Task-label perturbations further reduce system separation, confirming the practical relevance of task-adaptive scoring. These results support CalibratedRubric as an efficient, uncertainty-aware approach to open-ended LLM evaluation, with calibration gains depending on sufficient judge redundancy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。