arXiv:2501.06582cs.CL2025-01ACL被引 17

首个专家标注的法律合同条款检索数据集,助力AI辅助合同起草。

ACORD: An Expert-Annotated Retrieval Dataset for Legal Contract Drafting

  • 构建专家标注的合同条款检索数据集,聚焦复杂条款。
  • 包含114个查询和超12.6万对查询-条款,评分1~5星。
  • 适合法律AI研究者与合同自动化工具开发者使用。

信息检索,尤其是合同条款检索,是合同起草的基础,因为律师通常不会从零开始撰写合同,而是查找并修改最相关的先例条款。我们提出了Atticus Clause Retrieval Dataset(ACORD),这是首个由专家完整标注的合同起草检索基准数据集。ACORD聚焦于复杂合同条款,如责任限制、赔偿条款、控制权变更和最优待遇等。数据集包含114个查询和超过126,000个查询-条款对,每对均按1至5星进行评分。任务目标是为给定查询找到最相关的先例条款。基于双编码器检索器与点式大语言模型重排序器的组合表现出良好效果,但仍需显著提升以应对律师日常处理的复杂法律任务。作为首个由专家标注的合同起草检索基准,ACORD可为自然语言处理社区提供重要参考。

原文摘要 · Abstract (English)

Information retrieval, specifically contract clause retrieval, is foundational to contract drafting because lawyers rarely draft contracts from scratch; instead, they locate and revise the most relevant precedent. We introduce the Atticus Clause Retrieval Dataset (ACORD), the first retrieval benchmark for contract drafting fully annotated by experts. ACORD focuses on complex contract clauses such as Limitation of Liability, Indemnification, Change of Control, and Most Favored Nation. It includes 114 queries and over 126,000 query-clause pairs, each ranked on a scale from 1 to 5 stars. The task is to find the most relevant precedent clauses to a query. The bi-encoder retriever paired with pointwise LLMs re-rankers shows promising results. However, substantial improvements are still needed to effectively manage the complex legal work typically undertaken by lawyers. As the first retrieval benchmark for contract drafting annotated by experts, ACORD can serve as a valuable IR benchmark for the NLP community.

法律AI信息检索合同生成数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。