构建跨法域合同条款等效性数据集,助力跨国法律文本理解
LAUKIN: A Multi-jurisdictional Common Law Contract Dataset

- 基于多阶段检索与重排序构建条款配对,专家标注3000对等效性
- 包含14727对条款,最佳模型宏F1达65.11%,挑战显著
- 适用于法律NLP、跨法域合同分析及半监督学习研究
跨国公司日益需要跨法域合同审查,但现有法律NLP数据集多局限于单一法域。本文提出LAUKIN(澳大利亚、英国、印度法律等效性数据集),包含澳-英、英-印、印-澳三组条款对,标注布尔等效性。通过多阶段检索与重排序管道构建初始配对,再由法律专家标注3000对(900训练、600验证、1500测试)。数据集涵盖204份合同、8类协议,共14727对条款。评估12种模型,最优宏F1为65.11%,表明尽管共享法律传统,但起草惯例差异大,跨法域等效性判断复杂。此外,11727对未标注数据支持未来半监督学习研究。
原文摘要 · Abstract (English)
Multinational companies increasingly require cross-jurisdictional contract review, yet existing legal NLP datasets are largely restricted to a single jurisdiction. We introduce LAUKIN (Legal equivalence dataset of Australia, UK, and INdia), a dataset of clause pairs (AU-UK, UK-IN, IN-AU) labelled for boolean legal equivalence. We develop a novel multi-stage retrieval and reranking pipeline to construct the initial clause pair mapping, with a subset of clause pairs subsequently annotated by legal experts as Equivalent or Not Equivalent. The dataset comprises 14,727 clause pairs from 204 contracts across 8 agreement types, of which 3,000 are manually labelled: 900 train, 600 dev, and 1,500 test. We evaluate 12 models across 4 techniques, achieving a best macro-F1 of 65.11%, establishing LAUKIN as a challenging benchmark. Results reveal that, despite shared legal heritage, drafting conventions diverge significantly across jurisdictions, making cross-jurisdictional equivalence classification non-trivial. LAUKIN also includes 11,727 unlabelled training pairs to support future semi-supervised learning research in legal NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。