构建首个法罗语抽取式问答数据集,助力小语种NLP研究。
FoQA: A Faroese Question-Answering Dataset
- 用大模型生成+母语者验证,半自动化构建法罗语QA数据。
- 含2000个高质量样本,支持多模型在法罗语上的性能评估。
- 适合小语种自然语言处理、低资源语言研究者使用。
我们提出了FoQA,一个包含2,000个样本的法罗语抽取式问答(QA)数据集,采用半自动化方法结合大型语言模型(LLMs)与人工验证构建。数据源自法罗语维基百科文章,首先使用GPT-4-turbo生成初始问答对,随后通过问题重述提升复杂度,并由母语者进行质量验证。我们提供了多个模型(包括大模型和BERT)在FoQA上的基准性能指标,证明其在评估法罗语问答能力方面的有效性。数据集以三种版本发布:经验证的2,000个样本集、全部生成的10,001个样本集,以及用于错误分析的2,395个被拒样本集。
原文摘要 · Abstract (English)
We present FoQA, a Faroese extractive question-answering (QA) dataset with 2,000 samples, created using a semi-automated approach combining Large Language Models (LLMs) and human validation. The dataset was generated from Faroese Wikipedia articles using GPT-4-turbo for initial QA generation, followed by question rephrasing to increase complexity and native speaker validation to ensure quality. We provide baseline performance metrics for FoQA across multiple models, including LLMs and BERT, demonstrating its effectiveness in evaluating Faroese QA performance. The dataset is released in three versions: a validated set of 2,000 samples, a complete set of all 10,001 generated samples, and a set of 2,395 rejected samples for error analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。