arXiv:2510.19334stat.MLcs.AI2025-10被引 2

用大模型自动提取合同关键条款,提升法律审查效率。

Metadata Extraction Leveraging Large Language Models

  • 结合提示工程与分块策略,优化法律文本解析
  • 在CUAD等数据集上实现高精度条款识别
  • 适合法律科技公司和企业法务快速部署

大型语言模型的兴起彻底改变了多个领域的任务,包括法律文件分析自动化,这是现代合同管理系统的关键组成部分。本文提出了一种针对合同审查的增强型元数据提取综合方案,重点实现对重要法律条款的自动检测与标注。基于公开的Contract Understanding Atticus Dataset(CUAD)及专有合同数据集,研究展示了先进LLM方法与实际应用的融合。我们识别出三个优化元数据提取的关键要素:稳健的文本转换、策略性分块选择,以及链式思维(CoT)提示和结构化工具调用等高级LLM技术。实验结果表明,条款识别的准确性和效率均有显著提升。该方法有望大幅降低合同审查的时间与成本,同时保持高精度的法律条款识别能力。研究显示,经过精心优化的LLM系统可为法律专业人士提供有力支持,使各类规模组织更便捷地获得高效的合同审查服务。

原文摘要 · Abstract (English)

The advent of Large Language Models has revolutionized tasks across domains, including the automation of legal document analysis, a critical component of modern contract management systems. This paper presents a comprehensive implementation of LLM-enhanced metadata extraction for contract review, focusing on the automatic detection and annotation of salient legal clauses. Leveraging both the publicly available Contract Understanding Atticus Dataset (CUAD) and proprietary contract datasets, our work demonstrates the integration of advanced LLM methodologies with practical applications. We identify three pivotal elements for optimizing metadata extraction: robust text conversion, strategic chunk selection, and advanced LLM-specific techniques, including Chain of Thought (CoT) prompting and structured tool calling. The results from our experiments highlight the substantial improvements in clause identification accuracy and efficiency. Our approach shows promise in reducing the time and cost associated with contract review while maintaining high accuracy in legal clause identification. The results suggest that carefully optimized LLM systems could serve as valuable tools for legal professionals, potentially increasing access to efficient contract review services for organizations of all sizes.

大模型合同分析元数据提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。