arXiv:2501.16011cs.CL2025-01被引 1

专为西班牙语法律文本打造的高效语言模型

MEL: Legal Spanish Language Model

  • 基于XLM-RoBERTa-large,用西班牙官方公报等法律文献微调
  • 在法律文本理解任务中显著优于基线模型
  • 适合法律AI、司法科技领域研究者使用

法律文本以复杂专业的术语著称,处理难度大。加入西班牙语这类低资源语言后挑战更甚。尽管XLM-RoBERTa等预训练模型具备多语言能力,但在特定领域文档上的表现仍待探索。本文提出MEL——一个基于XLM-RoBERTa-large的法律语言模型,通过在西班牙官方公报(BOE)和议会文件等法律文本上进行微调,完成了数据收集、处理、训练与评估全过程。评估结果显示,该模型在理解西班牙语法律语言方面显著优于基线模型。案例研究进一步展示了其在新法律文本上的应用潜力,可在多种自然语言处理任务中取得优异表现。

原文摘要 · Abstract (English)

Legal texts, characterized by complex and specialized terminology, present a significant challenge for Language Models. Adding an underrepresented language, such as Spanish, to the mix makes it even more challenging. While pre-trained models like XLM-RoBERTa have shown capabilities in handling multilingual corpora, their performance on domain specific documents remains underexplored. This paper presents the development and evaluation of MEL, a legal language model based on XLM-RoBERTa-large, fine-tuned on legal documents such as BOE (Boletín Oficial del Estado, the Spanish oficial report of laws) and congress texts. We detail the data collection, processing, training, and evaluation processes. Evaluation benchmarks show a significant improvement over baseline models in understanding the legal Spanish language. We also present case studies demonstrating the model's application to new legal texts, highlighting its potential to perform top results over different NLP tasks.

法律AI西班牙语语言模型XLM-R

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。