arXiv:2506.15978cs.CLcs.AI2025-06被引 1

构建首个越南语文本分段与多选阅读理解数据集,助力低资源语言NLP研究

A Vietnamese Dataset for Text Segmentation and Multiple Choices Reading Comprehension

  • 基于维基百科构建越南语分段与多选题数据集
  • mBERT在阅读理解上达88.01%准确率,分段F1为63.15%
  • 验证多语言模型对低资源语言的有效性,适合多语言NLP研究者

越南语作为全球第20大语言,拥有超过1亿母语使用者,但在文本分段和机器阅读理解等关键自然语言处理任务上仍缺乏可靠资源。为此,我们提出了VSMRC——越南语文本分段与多选阅读理解数据集。该数据集源自越南语维基百科,包含15,942篇文档用于文本分段,以及16,347对经人工质量保证生成的合成多选问答对,确保数据可靠且多样。实验表明,mBERT在两项任务中均优于单语模型,在阅读理解测试集上达到88.01%准确率,文本分段测试集F1得分为63.15%。分析显示,多语言模型在越南语任务中表现更优,提示其在其他低资源语言中的潜在应用价值。VSMRC已公开于HuggingFace。

原文摘要 · Abstract (English)

Vietnamese, the 20th most spoken language with over 102 million native speakers, lacks robust resources for key natural language processing tasks such as text segmentation and machine reading comprehension (MRC). To address this gap, we present VSMRC, the Vietnamese Text Segmentation and Multiple-Choice Reading Comprehension Dataset. Sourced from Vietnamese Wikipedia, our dataset includes 15,942 documents for text segmentation and 16,347 synthetic multiple-choice question-answer pairs generated with human quality assurance, ensuring a reliable and diverse resource. Experiments show that mBERT consistently outperforms monolingual models on both tasks, achieving an accuracy of 88.01% on MRC test set and an F1 score of 63.15\% on text segmentation test set. Our analysis reveals that multilingual models excel in NLP tasks for Vietnamese, suggesting potential applications to other under-resourced languages. VSMRC is available at HuggingFace

越南语阅读理解文本分段低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。