arXiv:2412.04119cs.CL2024-12ACL被引 3

构建罗马尼亚法律问答数据集与知识图谱,提升低资源语言法律问答性能

GRAF: Graph Retrieval Augmented by Facts for Romanian Legal Multi-Choice Question Answering

  • 基于法律文本构建知识图谱,结合检索增强生成解决法律多选题
  • 在10,836个问题上达到接近主流方法的准确率,部分场景超越现有模型
  • 适合法律AI研究者和低资源语言NLP开发者参考

预训练语言模型在近年展现出卓越性能,推动了自然语言处理的新范式。法律领域因其文本特性受到关注,其中问答任务尤为突出。本文针对低资源语言罗马尼亚语,探索法律多选题问答任务。贡献包括:首次公开发布包含三个考试、共10,836个问题的罗马尼亚法律多选题数据集JuRO;构建包含93个法规文件及其763个时间跨度修改记录的法律语料库CROL,用于信息检索;首次提出罗马尼亚语法律知识图谱Law-RoG;并提出新型问答方法GRAF,该方法在多数设置下表现优于或媲美主流SOTA模型。

原文摘要 · Abstract (English)

Pre-trained Language Models (PLMs) have shown remarkable performances in recent years, setting a new paradigm for NLP research and industry. The legal domain has received some attention from the NLP community partly due to its textual nature. Some tasks from this domain are represented by question-answering (QA) tasks. This work explores the legal domain Multiple-Choice QA (MCQA) for a low-resource language. The contribution of this work is multi-fold. We first introduce JuRO, the first openly available Romanian legal MCQA dataset, comprising three different examinations and a number of 10,836 total questions. Along with this dataset, we introduce CROL, an organized corpus of laws that has a total of 93 distinct documents with their modifications from 763 time spans, that we leveraged in this work for Information Retrieval (IR) techniques. Moreover, we are the first to propose Law-RoG, a Knowledge Graph (KG) for the Romanian language, and this KG is derived from the aforementioned corpus. Lastly, we propose a novel approach for MCQA, Graph Retrieval Augmented by Facts (GRAF), which achieves competitive results with generally accepted SOTA methods and even exceeds them in most settings.

法律AI知识图谱多选题低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。