为大模型评判标准设计可衡量的规范,提升评估可靠性。
PReMISE: Policy Rubrics as Measurement Specifications for LLM Judges

- 将评判标准视为测量规范,通过数据发现最优政策级评分规则集。
- 发现现有标准均无法同时满足可靠、符合偏好、抗攻击三要素,而新框架在三方面表现均衡。
- 提出两种修复策略:优先级筛选提升判别准确率至68.6%,约束优化降低恶意响应得分率至36.0%。
大语言模型(LLM)裁判广泛用于评估开放性回答,但其评分高度依赖所用评判标准。模糊标准如“有用且事实正确”可能奖励精心包装却虚构事实或违背用户意图的回答。本文将可复用的评判标准视为测量规范:改变标准即改变由固定裁判生成的质量评估结果。我们提出PReMISE框架,基于成对人类偏好数据,(i) 发现政策层级的评判标准集合,(ii) 沿四个维度审计任意标准集:结构合理性、可靠性、偏好拟合度与对抗鲁棒性。结果显示,不同来源的标准无一能同时满足可靠、预测偏好、抗攻击;高评者间一致性并不意味着低可利用性。PReMISE是唯一在适用性、具体性和有效维度上均取得非平凡表现的来源。我们提出两种面向审计的修复操作:偏好排序选择将裁判在成对响应上的准确率从65.0%提升至68.6%,媲美最强基线,在跨裁判测试中领先三个裁判中的两个;可靠性约束精炼将恶意响应获得高分的比例从46.4%降至36.0%,评者间一致性仅微降(α从0.531降至0.519)。
原文摘要 · Abstract (English)
LLM judges are increasingly used to evaluate open-ended responses, but their scores depend strongly on the rubrics that condition them. A vague rubric asking for a response to be ``helpful and factual'' can reward polished answers that invent facts or violate user intent. We treat reusable rubrics as measurement specifications: changing the rubric changes the response quality measurement induced by a fixed judge. We introduce PReMISE, a framework that, given pairwise human-preference data, (i) discovers a policy-level rubric set, and (ii) audits any rubric set under LLM-judge use along four axes: structural adequacy, reliability, preference fit, and adversarial robustness. Across rubric sources no raw source is simultaneously reliable, preference-predictive, and adversarially robust; and high inter-rater agreement does not imply low exploitability. PReMISE is the only rubric source to score non-trivially on applicability, specificity, and effective dimensionality simultaneously. We contribute two audit-targeted repair operations: preference-rank selection raises judge accuracy on paired responses from $65.0\%$ to $68.6\%$, competitive with the strongest rubric-discovery baselines and leading on two of three judges in our cross-judge sweep; reliability-constrained refinement reduces the rate at which exploit responses receive high scores from $46.4\%$ to $36.0\%$ with little change in inter-judge agreement ($α{=}.531\to.519$).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。