arXiv:2511.00340cs.AI2025-11Conference of the …被引 7

首个专用于检测大模型法律推理缺陷的基准测试

Better Call CLAUSE: A Discrepancy Benchmark for Auditing LLMs Legal Reasoning Capabilities

  • 基于真实合同数据生成7500+扰动样本,设计10类法律异常
  • 发现主流大模型对细微法律错误识别率低,解释能力更弱
  • 适合法律AI安全评估、合规系统开发人员使用

大语言模型在高风险法律场景中的应用日益广泛,但缺乏系统性评测其法律推理可靠性的基准。为此,我们提出首个针对法律推理脆弱性的基准CLAUSE。通过从CUAD和ContractNLI等基础数据集生成超过7500份真实世界扰动合同,构建了10种不同类型的异常类别。采用角色驱动的生成流程,并利用检索增强生成(RAG)系统与官方法规比对,确保法律准确性。使用CLAUSE评估主流大模型在检测嵌入式法律缺陷及其解释能力上的表现。结果表明,这些模型常遗漏细微错误,且难以进行合法化解释。本研究为识别并修正法律AI中的推理缺陷提供了路径。

原文摘要 · Abstract (English)

The rapid integration of large language models (LLMs) into high-stakes legal work has exposed a critical gap: no benchmark exists to systematically stress-test their reliability against the nuanced, adversarial, and often subtle flaws present in real-world contracts. To address this, we introduce CLAUSE, a first-of-its-kind benchmark designed to evaluate the fragility of an LLM's legal reasoning. We study the capabilities of LLMs to detect and reason about fine-grained discrepancies by producing over 7500 real-world perturbed contracts from foundational datasets like CUAD and ContractNLI. Our novel, persona-driven pipeline generates 10 distinct anomaly categories, which are then validated against official statutes using a Retrieval-Augmented Generation (RAG) system to ensure legal fidelity. We use CLAUSE to evaluate leading LLMs' ability to detect embedded legal flaws and explain their significance. Our analysis shows a key weakness: these models often miss subtle errors and struggle even more to justify them legally. Our work outlines a path to identify and correct such reasoning failures in legal AI.

法律AI基准测试大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。