arXiv:2504.02864cs.CL2025-04

百万份美股公司合同数据集,助力法律AI研究与分析

The Material Contracts Corpus

  • 用LLaMA-2等NLP技术对100万份合同分类并关联主体
  • 发现雇佣与证券协议占SEC文件超八成,语言随时间变复杂
  • 支持批量下载,适合法律科技、金融合规领域研究者使用

本文发布公开可获取的《材料合同语料库》(Material Contracts Corpus, MCC),包含2000至2023年间美国证券交易委员会(SEC)收录的超过一百万份上市公司合同。该语料库支持合同设计与法律语言的实证研究,并推动基于AI的法律工具开发。通过机器学习与自然语言处理技术,包括微调的LLaMA-2模型,对合同按协议类型分类并关联具体签署方。语料库提供备案表格、文档格式及修订状态等元数据。研究揭示了合同语言、长度与复杂度随时间演变的趋势,指出雇佣协议与证券协议在申报文件中占主导地位。该资源可通过 https://mcc.law.stanford.edu 批量下载或在线访问。

原文摘要 · Abstract (English)

This paper introduces the Material Contracts Corpus (MCC), a publicly available dataset comprising over one million contracts filed by public companies with the U.S. Securities and Exchange Commission (SEC) between 2000 and 2023. The MCC facilitates empirical research on contract design and legal language, and supports the development of AI-based legal tools. Contracts in the corpus are categorized by agreement type and linked to specific parties using machine learning and natural language processing techniques, including a fine-tuned LLaMA-2 model for contract classification. The MCC further provides metadata such as filing form, document format, and amendment status. We document trends in contractual language, length, and complexity over time, and highlight the dominance of employment and security agreements in SEC filings. This resource is available for bulk download and online access at https://mcc.law.stanford.edu.

法律AI合同分析大数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。