arXiv:2608.20204cs.AIcs.CL2026-08综述

首个法律合同终审评估基准,测试大模型纠错能力

ContractScrub: A benchmark for final review of legal contracts

论文配图:ContractScrub: A benchmark for final review of legal contracts
图 1 · 摘自论文原文
  • 构建由律师手工设计的合同错误数据集,覆盖术语误用等三类问题
  • 顶尖模型宏平均召回率仅0.75,远低于通用基准表现
  • 适合法律AI研究者、合规系统开发者参考

法律工作高度依赖长文本处理,是大语言模型应用的重要领域。合同'清洗'——即对交易协议进行最终审查以发现错误与不一致——特别适合自动化,因其流程重复、耗时耗力且需细致阅读长文档。该任务与前沿大模型在长上下文推理、一致性检测和命名实体识别方面的能力天然契合。尽管具有显著经济价值和自动化潜力,此前尚未有针对合同清洗任务的正式评估。本文提出ContractScrub,首个专用于评估合同清洗能力的基准,包含由经验律师手工构建的合同,涵盖定义术语误用、引用错误、语言不一致等多样错误类型。结果显示,当前前沿模型表现令人意外地差,仅一个模型达到0.75的宏平均召回率,即使其在相关通用基准上表现优异,揭示了现有模型的实际局限性,强调了针对特定领域设计评估基准的重要性。

原文摘要 · Abstract (English)

Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs. Contract ``scrubbing,'' the final review of transactional agreements for errors and inconsistencies, is a particularly suitable task for automation, because it is routine, painstaking work requiring detailed attention to long documents. Scrubbing also seems to align naturally with the general capabilities expected of frontier LLMs around long-context reasoning, consistency checking, and named entity recognition (NER). Despite the economic value and potential for automation, no formal evaluations of LLMs performing contract scrubbing have been conducted. We introduce ContractScrub, the first benchmark designed to evaluate contract scrubbing capabilities, comprising contracts hand-crafted by experienced lawyers over diverse error categories such as misuse of defined terms, incorrect references, and inconsistent language. Frontier models perform surprisingly poorly with only one model reaching 0.75 macro average recall despite strong performance on seemingly related general benchmarks, demonstrating the practical limits of current models and the importance of narrowly targeted, domain-specific benchmarks for measuring real-world impact.

法律AI大模型评测合同分析自然语言推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。