arXiv:2410.13013cs.CLcs.AI2024-10被引 3

首个乌尔都语法律问答数据集,助力巴基斯坦宪法智能查询

LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering

  • 从巴基斯坦宪法提取中英双语法律问答对,人工与GPT-4协同构建
  • Claude-3.5-Sonnet在人工评估下达99.19%准确率,表现最佳
  • 专为低资源语言法律NLP设计,适合法律智能系统开发者使用

我们提出LEGAL-UQA,首个基于巴基斯坦宪法的乌尔都语-英语法律问答数据集。该平行数据集包含619组问答对,每对均附有对应的法律条文上下文,填补了低资源语言领域专用NLP资源的空白。数据集构建过程包括光学字符识别(OCR)提取、人工修正以及利用GPT-4辅助翻译与问答对生成。实验评估了当前主流通用语言模型与嵌入模型在LEGAL-UQA上的表现,其中Claude-3.5-Sonnet在人工评估中达到99.19%的准确率。我们还微调了mt5-large-UQA-1.0模型,揭示了多语言模型向专业领域迁移的挑战。此外,检索性能评估显示OpenAI的text-embedding-3-large优于Mistral的mistral-embed。LEGAL-UQA连接了全球NLP进展与本地化应用,尤其在宪法法律领域,为提升巴基斯坦的法律信息获取能力奠定基础。

原文摘要 · Abstract (English)

We present LEGAL-UQA, the first Urdu legal question-answering dataset derived from Pakistan's constitution. This parallel English-Urdu dataset includes 619 question-answer pairs, each with corresponding legal article contexts, addressing the need for domain-specific NLP resources in low-resource languages. We describe the dataset creation process, including OCR extraction, manual refinement, and GPT-4-assisted translation and generation of QA pairs. Our experiments evaluate the latest generalist language and embedding models on LEGAL-UQA, with Claude-3.5-Sonnet achieving 99.19% human-evaluated accuracy. We fine-tune mt5-large-UQA-1.0, highlighting the challenges of adapting multilingual models to specialized domains. Additionally, we assess retrieval performance, finding OpenAI's text-embedding-3-large outperforms Mistral's mistral-embed. LEGAL-UQA bridges the gap between global NLP advancements and localized applications, particularly in constitutional law, and lays the foundation for improved legal information access in Pakistan.

法律问答低资源语言乌尔都语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。