arXiv:2508.03080cs.AI2025-08被引 9

首个评估开源大模型识别合同法律风险能力的基准,发现其仍逊于商用模型。

ContractEval: Benchmarking LLMs for Clause-Level Legal Risk Identification in Commercial Contracts

  • 基于CUAD数据集对比4种商用与15种开源模型的条款级风险识别能力。
  • 商用模型在准确性和输出效果上整体领先,部分开源模型在特定维度表现接近。
  • 开源模型更易出现'无相关条款'误判,适合需高可靠性的法律场景研究者参考。

大型语言模型(LLMs)在法律等专业领域的潜力尚未充分探索。针对本地部署开源模型以保护数据隐私的需求,本文提出ContractEval,首个全面评估开源LLMs在商业合同中识别条款级法律风险的能力的基准。基于Contract Understanding Atticus Dataset(CUAD),我们评估了4种商用模型和15种开源模型。结果揭示五个关键发现:(1)商用模型在正确性和输出有效性上整体优于开源模型,尽管部分开源模型在特定维度表现接近;(2)更大规模的开源模型通常表现更好,但性能提升随模型增大而放缓;(3)推理模式提升了输出有效性,但降低了正确性,可能因对简单任务过度复杂化;(4)即使存在相关条款,开源模型也更频繁生成“无相关条款”响应,暗示思维惰性或提取信心不足;(5)模型量化虽加速推理,但导致性能下降,体现效率与精度的权衡。结果表明,多数模型表现相当于初级法律助理水平,开源模型需针对性微调以确保高风险法律场景中的准确性与有效性。ContractEval为未来法律领域大模型研发提供了坚实基准。

原文摘要 · Abstract (English)

The potential of large language models (LLMs) in specialized domains such as legal risk analysis remains underexplored. In response to growing interest in locally deploying open-source LLMs for legal tasks while preserving data confidentiality, this paper introduces ContractEval, the first benchmark to thoroughly evaluate whether open-source LLMs could match proprietary LLMs in identifying clause-level legal risks in commercial contracts. Using the Contract Understanding Atticus Dataset (CUAD), we assess 4 proprietary and 15 open-source LLMs. Our results highlight five key findings: (1) Proprietary models outperform open-source models in both correctness and output effectiveness, though some open-source models are competitive in certain specific dimensions. (2) Larger open-source models generally perform better, though the improvement slows down as models get bigger. (3) Reasoning ("thinking") mode improves output effectiveness but reduces correctness, likely due to over-complicating simpler tasks. (4) Open-source models generate "no related clause" responses more frequently even when relevant clauses are present. This suggests "laziness" in thinking or low confidence in extracting relevant content. (5) Model quantization speeds up inference but at the cost of performance drop, showing the tradeoff between efficiency and accuracy. These findings suggest that while most LLMs perform at a level comparable to junior legal assistants, open-source models require targeted fine-tuning to ensure correctness and effectiveness in high-stakes legal settings. ContractEval offers a solid benchmark to guide future development of legal-domain LLMs.

法律AI大模型评测合同分析开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。