首个高质量越南语法律问答数据集,填补低资源语言法律NLP空白
VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering
- 构建专用于越南法律领域的大型标注数据集
- 涵盖法律信息检索与问答任务,支持主流模型评估
- 助力越南语法律AI研究,适合法律NLP与低资源语言研究者
大语言模型(LLMs)在多个领域取得显著进展,包括法律文本处理。将LLMs应用于法律任务是自然演进且日益重要。然而,其实际能力常被夸大。尽管如此,我们距离完全自动化法律任务仍很遥远。此外,法律体系具有高度领域特异性,各国及语言间差异显著。因此,为不同自然语言构建法律文本处理应用的需求迫切而巨大。但对越南语等低资源语言而言,面临资源与标注数据匮乏的挑战。监督训练、验证和微调所需的标注法律语料库至关重要。本文提出VLQA数据集,这是一个全面且高质量的越南语法律领域资源。我们还对数据集进行统计分析,并通过前沿模型在法律信息检索与问答任务上的实验,验证其有效性。
原文摘要 · Abstract (English)
The advent of large language models (LLMs) has led to significant achievements in various domains, including legal text processing. Leveraging LLMs for legal tasks is a natural evolution and an increasingly compelling choice. However, their capabilities are often portrayed as greater than they truly are. Despite the progress, we are still far from the ultimate goal of fully automating legal tasks using artificial intelligence (AI) and natural language processing (NLP). Moreover, legal systems are deeply domain-specific and exhibit substantial variation across different countries and languages. The need for building legal text processing applications for different natural languages is, therefore, large and urgent. However, there is a big challenge for legal NLP in low-resource languages such as Vietnamese due to the scarcity of resources and annotated data. The need for labeled legal corpora for supervised training, validation, and supervised fine-tuning is critical. In this paper, we introduce the VLQA dataset, a comprehensive and high-quality resource tailored for the Vietnamese legal domain. We also conduct a comprehensive statistical analysis of the dataset and evaluate its effectiveness through experiments with state-of-the-art models on legal information retrieval and question-answering tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。