构建荷兰语法律条文检索数据集,提升多语言法律信息获取效率
Bilingual BSARD: Extending Statutory Article Retrieval to Dutch
- 构建法荷双语法律条文平行数据集bBSARD,支持跨语言检索
- 零样本密集模型表现不如经典BM25,但微调小模型可媲美商业模型
- 适合法律AI、多语言信息检索研究者使用
法律条文检索对提升公众和专业人士获取法律信息的便利性至关重要。比利时等多语言国家因需处理多种语言的法律问题而面临独特挑战。在法语版比利时法律条文检索数据集(BSARD)基础上,我们推出其双语版本bBSARD,包含法语与荷兰语的平行法律条文,以及原始问题及其荷兰语翻译。基于bBSARD,我们在荷兰语和法语上对多种检索模型进行了广泛基准测试,涵盖词法模型、零样本密集模型及微调的小型基础模型。实验表明,尽管商业模型在零样本设置中表现更优,但通过微调小型语言专用模型,可达到甚至超越其性能;同时,BM25在两种语言中仍保持竞争力。本研究数据集与评估代码已公开。
原文摘要 · Abstract (English)
Statutory article retrieval plays a crucial role in making legal information more accessible to both laypeople and legal professionals. Multilingual countries like Belgium present unique challenges for retrieval models due to the need for handling legal issues in multiple languages. Building on the Belgian Statutory Article Retrieval Dataset (BSARD) in French, we introduce the bilingual version of this dataset, bBSARD. The dataset contains parallel Belgian statutory articles in both French and Dutch, along with legal questions from BSARD and their Dutch translation. Using bBSARD, we conduct extensive benchmarking of retrieval models available for Dutch and French. Our benchmarking setup includes lexical models, zero-shot dense models, and fine-tuned small foundation models. Our experiments show that BM25 remains a competitive baseline compared to many zero-shot dense models in both languages. We also observe that while proprietary models outperform open alternatives in the zero-shot setting, they can be matched or surpassed by fine-tuning small language-specific models. Our dataset and evaluation code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。