构建超大规模西班牙语检索数据集,助力母语者信息获取
MessIRve: A Large-Scale Spanish Information Retrieval Dataset
- 从谷歌搜索补全抓取近70万条真实西班牙语查询
- 覆盖多地区方言,文档来自维基百科,主题广泛
- 为西班牙语信息检索研究提供基准数据集
信息检索(IR)是根据用户查询找出相关文档的任务。尽管西班牙语是第二大母语使用语言,但现有的西班牙语IR数据集较少,限制了针对西班牙语使用者的信息获取工具发展。我们提出了MessIRve,一个大规模西班牙语信息检索数据集,包含近70万条来自谷歌搜索补全API的查询,相关文档均来自维基百科。该数据集的查询涵盖多样化的西班牙语使用区域,不同于其他通过英语翻译或忽略方言差异构建的数据集。其大规模特性使其能够覆盖广泛的主题,超越现有小规模数据集的局限。我们提供了数据集的详细描述、与现有数据集的对比分析,以及主流IR模型的基线评估。本工作旨在推动西班牙语信息检索研究,提升西班牙语用户的资讯可及性。
原文摘要 · Abstract (English)
Information retrieval (IR) is the task of finding relevant documents in response to a user query. Although Spanish is the second most spoken native language, there are few Spanish IR datasets, which limits the development of information access tools for Spanish speakers. We introduce MessIRve, a large-scale Spanish IR dataset with almost 700,000 queries from Google's autocomplete API and relevant documents sourced from Wikipedia. MessIRve's queries reflect diverse Spanish-speaking regions, unlike other datasets that are translated from English or do not consider dialectal variations. The large size of the dataset allows it to cover a wide variety of topics, unlike smaller datasets. We provide a comprehensive description of the dataset, comparisons with existing datasets, and baseline evaluations of prominent IR models. Our contributions aim to advance Spanish IR research and improve information access for Spanish speakers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。