构建首个覆盖西亚北非10种语言的问答基准,聚焦真实信息查询场景。
TyDi QA-WANA: A Benchmark for Information-Seeking Question Answering in Languages of West Asia and North Africa
- 直接在10种语言原生语境收集数据,避免翻译带来的文化偏差。
- 含2.8万条问题-文章配对,文章篇幅大,考验模型长文本理解能力。
- 适用于评估多语言信息检索与跨语言问答模型,推动低资源语言研究。
我们提出TyDi QA-WANA,一个包含28,000个样本的问答数据集,涵盖西亚和北非地区的10种语言变体。数据收集旨在激发真实的信息寻求型问题,即提问者确实希望获取答案。每个问题均配有一篇可能包含答案的完整文章,文章规模较大,适合评估模型在处理长文本上下文时的答题能力。所有数据均在各语言原生语境中直接采集,未使用翻译,以确保文化相关性。我们报告了两个基线模型的表现,并公开代码与数据,以促进研究社区进一步改进。
原文摘要 · Abstract (English)
We present TyDi QA-WANA, a question-answering dataset consisting of 28K examples divided among 10 language varieties of western Asia and northern Africa. The data collection process was designed to elicit information-seeking questions, where the asker is genuinely curious to know the answer. Each question in paired with an entire article that may or may not contain the answer; the relatively large size of the articles results in a task suitable for evaluating models' abilities to utilize large text contexts in answering questions. Furthermore, the data was collected directly in each language variety, without the use of translation, in order to avoid issues of cultural relevance. We present performance of two baseline models, and release our code and data to facilitate further improvement by the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。