arXiv:2601.04424cs.CL2026-01中稿 · EMNLP

用检查清单+智能代理评估大模型在长篇法律摘要中的表现

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

  • 设计双模块评估框架:带检查清单的参考对比和无参考的智能代理
  • 模型更易遗漏关键信息而非编造,长文本下性能明显下降
  • 智能代理节省36%以上计算资源,跨领域通用性强

大型语言模型(LLMs)现已支持长达100万标记的上下文,但其在复杂长文本任务中的优劣尚不明确。本文聚焦多文档法律案件摘要任务,单个案件常超过10万标记。我们系统评估了12个前沿LLM,采用Gavel框架,包含基于参考的Gavel-Ref(含检查清单、残余事实与写作风格评估)和无参考的Gavel-Agent(直接从源文档评估事实覆盖)。结果表明,当前模型更倾向于遗漏关键信息而非虚构内容;对简单检查项(如立案日期)表现良好,但在罕见复杂项(如和解条款)上表现差,且随案件长度增加性能下降。为元评估Gavel,我们收集了160小时人工标注数据。Gavel-Agent相比端到端与分块方法,至少减少36%的标记使用量,同时保持竞争力;该方法在医疗领域也表现优异,至少减少77%的标记消耗。

原文摘要 · Abstract (English)

Large language models (LLMs) now support contexts of up to 1M tokens, but their strengths and weaknesses on complex long-context tasks remain unclear. To study this, we focus on multi-document legal case summarization, where a single case often spans many documents exceeding 100K tokens. We systematically evaluate 12 frontier LLMs with Gavel, which consists of Gavel-Ref, a reference-based evaluation framework with checklist, residual-fact, and writing-style evaluations, and Gavel-Agent, a reference-free agent for evaluating factual coverage directly from source documents. Our results show that current models are more prone to omitting key information than hallucinating. They all perform well on simple checklist items, such as filing date, but struggle with rare and complex items, such as settlements. Performance also declines as case length increases. To meta-evaluate Gavel, we collect 160 hours of human annotations. Gavel-Agent reduces token usage by at least 36% compared to end-to-end and chunk-by-chunk methods while achieving competitive performance. Gavel-Agent also generalizes to the medical domain, performing the best with at least 77% fewer tokens.

大模型评估法律AI长文本生成智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。