arXiv:2502.20936cs.CLcs.AI2025-02被引 10

构建了9600万条多语言问答对数据集,助力跨语言信息检索研究。

WebFAQ: A Multilingual Collection of Natural Q&A Datasets for Dense Retrieval

  • 从网页FAQ结构化数据中提取9600万条自然问答对,覆盖75种语言。
  • 用于微调模型后,在多个跨语言检索任务上实现显著性能提升。
  • 生成超1000对高质量双语语料,支持多语言翻译与对齐研究。

我们提出WebFAQ,一个基于schema.org FAQ标注的大规模开放域问答数据集。该数据集包含9600万条自然语言问答对,覆盖75种语言,其中4700万条(49%)为非英语样本。WebFAQ还构成了20个单语检索基准测试的基础,总计1120万条问答对(含590万条非英语)。所有数据经过精细化过滤和近似重复检测,确保高质量。为验证其有效性,我们使用这些问答对微调XLM-RoBERTa模型,结果在零样本设置下对其他多语言检索基准也展现出显著性能提升。此外,利用WebFAQ构建了跨越1000余种语言对的对齐双语语料库,采用先进的一体化双语句对挖掘与大模型评估的自动翻译质量检验方法,生成的语料库翻译质量优于同类数据集。WebFAQ及所有相关资源已开源,可在GitHub和HuggingFace获取。

原文摘要 · Abstract (English)

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from FAQ-style schema.org annotations. In total, the data collection consists of 96 million natural question-answer (QA) pairs across 75 languages, including 47 million (49%) non-English samples. WebFAQ further serves as the foundation for 20 monolingual retrieval benchmarks with a total size of 11.2 million QA pairs (5.9 million non-English). These datasets are carefully curated through refined filtering and near-duplicate detection, yielding high-quality resources for training and evaluating multilingual dense retrieval models. To empirically confirm WebFAQ's efficacy, we use the collected QAs to fine-tune an in-domain pretrained XLM-RoBERTa model. Through this process of dataset-specific fine-tuning, the model achieves significant retrieval performance gains, which generalize - beyond WebFAQ - to other multilingual retrieval benchmarks evaluated in zero-shot setting. Last but not least, we utilize WebFAQ to construct a set of QA-aligned bilingual corpora spanning over 1000 language pairs using state-of-the-art bitext mining and automated LLM-assessed translation evaluation. Due to our advanced, automated method of bitext dataset generation, the resulting bilingual corpora demonstrate higher translation quality compared to similar datasets. WebFAQ and all associated resources are publicly available on GitHub and HuggingFace.

多语言问答对检索语料库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。