arXiv:2410.20301cs.IR2024-10

WindTunnel通过保留数据社区结构,高效生成大语料代表性样本。

WindTunnel -- A Framework for Community Aware Sampling of Large Corpora

  • 基于社区结构设计采样策略,避免信息失真
  • 显著降低检索实验的计算成本,支持快速评估
  • 适合需要大规模语料测试的搜索与生成系统研发

在搜索或检索增强生成等全面的信息检索实验中,高计算成本是一个常见问题。这是因为评估检索算法需对整个语料库进行索引,而该规模远超实际评估中的(查询,结果)对数量。这一问题在大数据和神经检索场景下尤为突出,索引过程变得日益耗时且复杂。本文提出WindTunnel,一个由Yext开发的新框架,用于生成大型语料库的代表性样本,从而实现高效的端到端信息检索实验。通过保持数据集的社区结构,WindTunnel克服了现有采样方法的局限性,提供更准确的评估结果。

原文摘要 · Abstract (English)

Conducting comprehensive information retrieval experiments, such as in search or retrieval augmented generation, often comes with high computational costs. This is because evaluating a retrieval algorithm requires indexing the entire corpus, which is significantly larger than the set of (query, result) pairs under evaluation. This issue is especially pronounced in big data and neural retrieval, where indexing becomes increasingly time-consuming and complex. In this paper, we present WindTunnel, a novel framework developed at Yext to generate representative samples of large corpora, enabling efficient end-to-end information retrieval experiments. By preserving the community structure of the dataset, WindTunnel overcomes limitations in current sampling methods, providing more accurate evaluations.

信息检索采样方法大语料

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。