arXiv:2410.12852cs.CLcs.LG2024-10被引 2

基于希腊语法律文本训练出4个大模型,性能超越现有同类模型。

The Large Language Model GreekLegalRoBERTa

  • 用希腊法律与非法律文本训练四个大语言模型
  • 在命名实体识别和法律主题分类任务中表现更优
  • 适合低资源语言法律NLP研究者参考

我们开发了四个版本的GreekLegalRoBERTa,即在希腊语法律与非法律文本上训练的大型语言模型。实验表明,我们的模型在涉及希腊法律文档的两项任务——命名实体识别与多类别法律主题分类中,优于GreekLegalBERT、GreekLegalBERT-v2和GreekBERT。本工作被视为利用现代自然语言处理技术与方法,推动低资源语言(如希腊语)领域特定NLP任务研究的重要贡献。

原文摘要 · Abstract (English)

We develop four versions of GreekLegalRoBERTa, which are four large language models trained on Greek legal and nonlegal text. We show that our models surpass the performance of GreekLegalBERT, Greek- LegalBERT-v2, and GreekBERT in two tasks involving Greek legal documents: named entity recognition and multi-class legal topic classification. We view our work as a contribution to the study of domain-specific NLP tasks in low-resource languages, like Greek, using modern NLP techniques and methodologies.

法律NLP大模型低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。