arXiv:2410.10229cs.CLcs.AI2024-10中稿 · to LREC-COLING 202…被引 8

构建首个高质量孟加拉语开放域问答数据集,助力低资源语言NLP发展。

BanglaQuAD: A Bengali Open-domain Question Answering Dataset

  • 由母语者基于孟加拉维基百科人工构建问答对
  • 包含30,808对高质量问题与答案,覆盖本土话题与术语
  • 提供本地化标注工具,适合低资源语言研究者使用

孟加拉语是全球第七大常用语言,但在自然语言处理领域仍属低资源语言。开放域问答任务需同时理解问题与文本内容,极具挑战性。目前针对孟加拉语的问答研究极少,且多数数据集通过英译孟加拉语生成,导致句式不自然、存在噪声。此外,这些数据缺乏与孟加拉文化相关的主题与专业术语。本文提出BanglaQuAD,一个基于孟加拉维基百科文章、由母语者构建的孟加拉语问答数据集,共包含30,808个问题-答案对。同时,我们开发了一款可在本地机器上运行的标注工具,支持高效构建问答数据集。定性分析表明该数据集具有较高质量。

原文摘要 · Abstract (English)

Bengali is the seventh most spoken language on earth, yet considered a low-resource language in the field of natural language processing (NLP). Question answering over unstructured text is a challenging NLP task as it requires understanding both question and passage. Very few researchers attempted to perform question answering over Bengali (natively pronounced as Bangla) text. Typically, existing approaches construct the dataset by directly translating them from English to Bengali, which produces noisy and improper sentence structures. Furthermore, they lack topics and terminologies related to the Bengali language and people. This paper introduces BanglaQuAD, a Bengali question answering dataset, containing 30,808 question-answer pairs constructed from Bengali Wikipedia articles by native speakers. Additionally, we propose an annotation tool that facilitates question-answering dataset construction on a local machine. A qualitative analysis demonstrates the quality of our proposed dataset.

问答系统低资源语言孟加拉语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。