arXiv:2606.01629cs.CL2026-06

评测大模型当裁判评估长文本的可靠性,发现现有方法仍不稳定。

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

论文配图:Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
图 1 · 摘自论文原文
  • 构建跨场景长文本评估基准LongJudgeBench,覆盖多种真实任务
  • 实验显示当前大模型裁判在不同场景下表现不稳,可靠性差
  • 适合关注自动文本评估、大模型评判机制的研究者

随着大语言模型在长文本生成中的广泛应用,可靠评估长文本输出成为关键挑战。以大模型为裁判(LLM-as-a-judge)提供了可扩展的人类评估替代方案,但其在长文本评估中的可靠性尚未充分检验:现有元评估基准主要聚焦短文本。与短文本评估不同,长文本评估不仅涉及输出长度,更需对整体结构、任务相关覆盖度与深度、跨段落一致性及场景特定质量标准进行复杂文档级判断。本文提出LongJudgeBench,一个涵盖多样化真实场景和评估协议的综合性基准,用于评估大模型裁判在长文本上的表现。我们系统测试了多种基础模型和评估设置下的大模型裁判。结果揭示显著的可靠性差距:当前大模型裁判在不同场景中仍不稳定,尽管评分标准或参考答案有所帮助,但并非总能弥补不足。我们希望LongJudgeBench能推动更鲁棒、上下文感知且与人类对齐的大模型裁判方法研究。代码已开源:https://github.com/cjj826/LongJudgeBench。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly used for long-form generation, reliably evaluating long-form outputs has become a critical challenge. LLM-as-a-judge offers a scalable alternative to human evaluation, yet its reliability in long-form output evaluation remains underexamined: existing meta-evaluation benchmarks focus mainly on short-form outputs. Compared with short-form evaluation, long-form evaluation is not merely a matter of output length; it often requires judges to make more complex document-level assessments of overall organization, task-relevant coverage and depth, cross-section consistency, and scenario-specific quality criteria. In this work, we introduce LongJudgeBench, a comprehensive benchmark for evaluating LLM judges on long-form outputs across diverse real-world scenarios and judging protocols. We systematically evaluate a broad range of LLM judges, covering multiple base models and judging settings. Our results reveal a substantial reliability gap: current LLM judges remain unstable across scenarios, and rubrics or references are helpful but not always sufficient. We hope LongJudgeBench will support future research on more robust, context-aware, and human-aligned LLM-as-a-judge methods. Our code is available at https://github.com/cjj826/LongJudgeBench.

大模型评估长文本生成自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。