arXiv:2607.19322cs.CL2026-07

用两级评分框架提升开放生成的事实完整性评估可靠性

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

论文配图:Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
图 1 · 摘自论文原文
  • 先设计结构化评分标准,再转为机器可读的二元检查项
  • 在1813个真实图像关联问题上测试,最佳模型得分58.7%
  • 适合评估长文本生成事实完整性的研究者和开发者

开放生成的基于评分标准的评估面临表达力与可靠性之间的根本矛盾。构建准确的评分标准需刻画优质回答的空间结构:开放的答案集合、有序的过程及事实的重要性权重。而评分时,人类判断者在简单二元判断上比处理复杂结构更可靠。为此,本文提出两级元评分框架:作者阶段使用结构化元评分标准,评估阶段通过固定机械规则将其转化为可被大模型稳定评分的二元检查清单。该框架被具体实现为Gamut(多模态事实性基础评估),一个用于长文本生成事实完整性的基准。Gamut包含1,813个问题,源自10个不同领域的实际可穿戴设备图像,每个问题均配有由专家验证的证据支持评分标准。对14个前沿及开源模型的评估显示,Gamut具有真实挑战性(最高分58.7%来自Gemini 3.1 Pro)、高度区分性且对评判者选择不敏感。

原文摘要 · Abstract (English)

Rubric-based evaluation of open-ended generation faces a fundamental tension between expressiveness and reliability. Authoring a faithful rubric requires expressing the structure of the space of good answers: open-ended sets of acceptable options, ordered processes, and the relative importance of facts. Grading with the rubric requires a judge to score consistently, and judges are far more reliable on flat, binary checks than on rich structure. We resolve this tension with a two-level meta-rubric framework. A structured meta-rubric captures the grading criteria at authoring time, and fixed mechanical rules compile it into a flat checklist of binary, machine-gradable checks that an LLM judge scores reliably at evaluation time. We instantiate the framework as Gamut (Grounded Assessment of Multimodal Factuality), a benchmark for factual completeness in long-form generation. Gamut comprises 1,813 questions grounded in real wearable imagery across 10 diverse domains, each paired with an evidence-backed rubric verified by expert human annotators. Evaluating 14 frontier and open-weight models, we find Gamut genuinely challenging (best score 58.7% from Gemini 3.1 Pro), highly discriminative, and robust to the choice of judge.

事实性评估开放生成评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。