arXiv:2608.18303cs.AIcs.LG2026-08

让大模型自检评估,拆解判断依据,提升结果可解释性。

SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition

论文配图:SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition
图 1 · 摘自论文原文
  • 从模型错误案例中自动提炼子问题,无需额外标注或微调。
  • 在RewardBench上达到92.7%准确率,接近专家模型表现。
  • 提供逐维度投票证据,可追溯判断偏差和标签模糊根源。

LLM-as-judge评估将响应质量判断简化为单一的二选一偏好选择,无法识别导致偏好的具体质量维度,也难以区分模型错误与真实标签歧义。我们提出SESSE(Sketch, Expand, Sort, Summarize, Evaluate)框架,一种无需训练的方法,通过直接从评判模型自身的错误案例中挖掘结构化子问题,实现对整体判断的分解;该方法无需参考答案、任务专用评分标准或微调。在RewardBench(n=1,000)上,SESSE的表现接近思维链基线,并与经过微调的RISE-Judge-32B(92.7%)相当,同时保持完全无训练特性。每个维度的投票证据提供了可解释的审计轨迹,可用于诊断标签歧义与评判失败模式,这是单一整体输出所无法提供的。

原文摘要 · Abstract (English)

LLM-as-judge evaluation reduces response quality assessment to a single holistic A/B preference choice, providing no mechanism to isolate which quality dimensions drove the preference or distinguish model errors from genuine label ambiguity. We propose SESSE (Sketch, Expand, Sort, Summarize, Evaluate), a training-free framework that decomposes holistic judgment into structured sub-questions mined directly from the judge's own error cases; requiring no oracle responses, task-specific rubrics, or fine-tuning. On RewardBench (n=1,000), SESSE achieves near-parity with the chain-of-thought baseline and is competitive with RISE-Judge-32B (92.7%), a fine-tuned specialist, while remaining fully training-free. Per-criterion vote evidence provides an interpretable audit trail for diagnosing label ambiguity and judge failure modes unavailable from a single holistic output token.

模型评估可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。