arXiv:2607.15092cs.CL2026-07

仅用一个问题生成可自我验证的评分标准,避免无效或偏颇的评判。

Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence

论文配图:Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence
图 1 · 摘自论文原文
  • 从零开始生成评分标准,仅依赖问题和自动生成的对比回答对。
  • 在五个评测集上平均准确率最优,七项中有六项领先。
  • 适合需要高质量、无外部标注评分标准的研究者使用。

评分标准为训练和评估大语言模型提供了结构化、细粒度的信号。然而,构建可靠的特定查询评分标准仍具挑战。现有方法常依赖人工编写的评分标准、偏好数据或采样响应。直接从查询生成评分标准虽无需外部资源,但无法确保生成的标准具备判别能力。该标准可能无法区分答案质量、奖励非必要风格,或惩罚合理替代策略。本文提出 Rubrics on Trial,一种仅依赖查询的框架,从空集开始演化评分标准,无需外部标注或模型训练。其通过合成的条件化响应对获取监督信号,并在加入前验证每个候选评分标准,筛选出不具备判别力、过度具体或仅关注风格的标准。在五个偏好基准测试套件上的实验表明,Rubrics on Trial 效果显著,在平均准确率上表现最佳,且在七项评估中领先六项。

原文摘要 · Abstract (English)

Rubrics provide structured, fine-grained signals for training and evaluating large language models (LLMs). Yet reliable query-specific rubrics are difficult to construct. Existing approaches often derive supervision from human-written rubrics, preference data, or sampled responses. Direct query-to-rubric generation avoids these resources, but provides no explicit check that a plausible rubric is useful. Such a rubric may fail to distinguish answer quality, reward an optional style, or penalize a valid alternative strategy. We introduce Rubrics on Trial, a query-only framework that evolves a rubric set from an empty set without external annotations or model training. It derives supervision solely from synthetic rubric-conditioned response pairs and validates each proposed rubric before adding it, screening out non-discriminative, over-specific, and style-only candidate rubrics. Experiments across five preference benchmark suites demonstrate the effectiveness of Rubrics on Trial, which achieves the best average accuracy and leads on six of seven evaluation sets.

评分标准大模型评估自生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。