用规则评分取代整体打分,让大模型评估更公平可靠
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese

- 采用多评委过滤的二元规则评分,降低评价偏差
- 在16个模型中,规则评分使评委一致率达100%,整体评分仅70%
- 评分差异扩大47%,提升模型区分度,适合开放任务评估
基于整体模型打分的排名对评判模型的偏见敏感。我们证明,转向使用多评委过滤的二元规则评分可消除这种敏感性:判断过程的分解比评判模型本身更重要。为此,我们提出Prosa——首个真实用户多轮巴西葡萄牙语聊天基准,包含1,000条来自WildChat的对话,由三个模型族的三位评委对16个模型进行评分。经规则评分过滤后,三位评委在全部16个模型的排名上达成一致;而整体评分下仅在7个排名上一致。此外,规则过滤流程使相邻模型间的平均得分差距提升47%,显著增强评测的区分能力。使用Gemini 3 Flash作为评委时,新模型在Prosa上的评估成本约为2.1美元。我们已公开该基准与过滤代码,确保未来模型可在相同条件下评估。该方法亦可复用于其他开放式评测场景。
原文摘要 · Abstract (English)
Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scoring with multi-judge filtering removes this sensitivity: decomposing the judgement matters more than the judge model itself. To support this claim, we introduce Prosa, the first real user multi-turn Brazilian Portuguese chat benchmark: 1,000 WildChat conversations scored by three judges from three model families on 16 models. Under filtered rubric scoring the three judges agree on every one of the 16 ranks, whereas under holistic scoring they agree on only 7 of 16. Additionally, the rubric filtering pipeline increases the average score gap between neighbouring models by 47%, thereby improving Prosa's discriminative power. Evaluating a new model on Prosa costs approximately $2.1 when using Gemini 3 Flash as the judge. We release the benchmark and the filtering code to ensure that future models can be assessed under identical conditions. These artifacts also make our rubric-based scoring method reusable beyond Prosa, supporting other open-ended evaluation settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。