评测大模型对多约束指令的逐项判断能力,发现高准确率不等于高稳定性。
MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following

- 设计了可细分判断的多约束评测基准MCJudgeBench,支持逐条约束评估。
- 发现模型在罕见标签(部分/无)上检测可靠性显著下降,且正确率与稳定性不正相关。
- 适合关注模型评估可靠性、安全性和细粒度判断能力的研究者使用。
多约束指令遵循要求验证响应是否满足多个独立要求,但现有大模型评判通常仅基于整体判断。本文提出MCJudgeBench,一个用于多约束指令遵循中约束级评判的基准。每个实例包含指令、候选响应、显式约束列表、每条约束的黄金标签(是/部分/否),以及受控的响应扰动。评估协议包含多种提示变体,用于测试评判者稳定性。我们使用正确率和不一致率指标,评估专有与开源大模型评判者,区分随机解码下的内在不一致与提示/响应扰动下的过程性不一致。结果表明,评判可靠性具有多维度特征:整体表现强并不意味着各类标签检测均可靠,尤其在罕见的‘部分’和‘否’类别中。高正确率模型未必具有低不一致率。增加推理环节可提升正确率,但不统一改善稳定性。这些发现强调需在约束级别评估大模型评判者,以识别其失效模式。
原文摘要 · Abstract (English)
Multi-constraint instruction following requires verifying whether a response satisfies multiple individual requirements, yet LLM judges are often assessed only through overall-response judgments. We introduce MCJudgeBench, a benchmark for constraint-level judge evaluation in multi-constraint instruction following. Each instance includes an instruction, a candidate response, an explicit constraint list, per-constraint gold labels in {yes, partial, no}, and controlled response-side perturbations. The evaluation protocol further includes evaluation prompt variants to test judge stability. We evaluate proprietary and open-source LLM judges using both correctness and inconsistency metrics, distinguishing intrinsic inconsistency under stochastic decoding from procedural inconsistency under prompt and response perturbations. Our results show that judge reliability has multiple dimensions: strong overall performance does not guarantee equally reliable detection across label categories, especially for rarer partial and no cases. Judges with higher correctness do not always have lower inconsistency. Evaluation with reasoning improves correctness but does not uniformly improve stability. These findings motivate evaluating LLM judges at the constraint level to study these failure modes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。