arXiv:2501.15990cs.CL2025-01被引 1

构建首个西班牙语合同条款法律语料库,助力法律NLP研究

3CEL: A corpus of legal Spanish contract clauses

  • 人工标注373份招标文件,涵盖19类合同关键信息
  • 共生成4782个标签,覆盖合同条款核心要素
  • 适合法律AI、司法文本分析领域研究人员使用

由于数据获取难与法律专业知识稀缺,西班牙语自然语言处理(NLP)领域的法律语料资源极为有限。INESData 2024是由马德里理工大学(UPM)牵头、知识工程研究所(IIC)开发的欧盟资助项目,旨在构建面向西班牙语法律/行政领域的先进NLP资源。本文介绍其中成果——西班牙语合同条款语料库(3CEL),该语料库基于人工标注的373份招标文件,采用19个定义类别(总计4782个标签),精准识别合同理解与审查所需的关键信息。

原文摘要 · Abstract (English)

Legal corpora for Natural Language Processing (NLP) are valuable and scarce resources in languages like Spanish due to two main reasons: data accessibility and legal expert knowledge availability. INESData 2024 is a European Union funded project lead by the Universidad Politécnica de Madrid (UPM) and developed by Instituto de Ingeniería del Conocimiento (IIC) to create a series of state-of-the-art NLP resources applied to the legal/administrative domain in Spanish. The goal of this paper is to present the Corpus of Legal Spanish Contract Clauses (3CEL), which is a contract information extraction corpus developed within the framework of INESData 2024. 3CEL contains 373 manually annotated tenders using 19 defined categories (4 782 total tags) that identify key information for contract understanding and reviewing.

法律NLP语料库合同分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。