arXiv:2501.11628cs.IR2025-01被引 13

测试稀疏检索算法在超大规模数据下的表现,发现其效率与效果随规模变化的关键规律。

Investigating the Scalability of Approximate Sparse Retrieval Algorithms to Massive Datasets

  • 对比Seismic和图结构方法在海量数据上的检索性能
  • 在1.38亿条数据上评估Splade模型,索引耗时显著上升
  • 揭示稀疏检索在超大规模场景下的可扩展性瓶颈

由于在顶k检索中表现优异且具有内在可解释性,学习得到的稀疏文本嵌入近年来受到关注。然而,其分布特性长期制约其在真实系统中的应用。近期出现的近似算法利用稀疏嵌入的分布特性加速检索,但现有研究多局限于仅数百万文档的数据集(如MSMARCO)。尚不清楚这些系统在更大规模下的行为及潜在挑战。为此,本文系统考察了当前先进检索算法在超大规模数据上的表现。对比分析了新提出的Seismic方法以及源自密集检索的图结构方案,并在包含1.38亿段落的MsMarco-v2数据集上,对Splade嵌入进行了广泛评估,报告了索引时间及其他效率与有效性指标。

原文摘要 · Abstract (English)

Learned sparse text embeddings have gained popularity due to their effectiveness in top-k retrieval and inherent interpretability. Their distributional idiosyncrasies, however, have long hindered their use in real-world retrieval systems. That changed with the recent development of approximate algorithms that leverage the distributional properties of sparse embeddings to speed up retrieval. Nonetheless, in much of the existing literature, evaluation has been limited to datasets with only a few million documents such as MSMARCO. It remains unclear how these systems behave on much larger datasets and what challenges lurk in larger scales. To bridge that gap, we investigate the behavior of state-of-the-art retrieval algorithms on massive datasets. We compare and contrast the recently-proposed Seismic and graph-based solutions adapted from dense retrieval. We extensively evaluate Splade embeddings of 138M passages from MsMarco-v2 and report indexing time and other efficiency and effectiveness metrics.

稀疏检索可扩展性大规模数据信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。