arXiv:2609.08860cs.IR2026-09

首个开源实现支持三种文档索引类型,解决生成式检索复现难题

REDSI: Addressing the Reproducibility and Evaluation Consistency of Differentiable Search Indexing for Document Retrieval

  • 提供完整开源实现,覆盖原子、朴素、语义三类文档标识符
  • 在NQ320K上达成与已有基线相当或更优的检索效果
  • 支持模型缩放实验,为未来研究提供可复现的评估框架

可微搜索索引(DSI)框架已成为生成式检索的事实基准。然而,DSI难以复现:缺乏公开实现覆盖全部三种原始文档标识符类型(原子、朴素、语义),报告结果差异大,且广泛使用的NQ320K数据集基于Natural Questions通过多种未明确定义的预处理构建。本文提出ReDSI,首个支持所有三种标识符类型的开源DSI实现,并提供可参数化、文档齐全的NQ320K构建流程。实验表明,其性能达到或超过先前的DSI基线。此外,我们在模型缩放条件下进行了大量实验,涵盖检索有效性、参数效率、训练方法和解码策略,为未来研究开辟新方向。

原文摘要 · Abstract (English)

The differentiable search index (DSI) framework (Tay et al., 2022) has become the de facto baseline for generative retrieval. However, DSI is hard to reproduce: no public implementation covers all three original document identifier types (atomic, naive, semantic), reported results vary widely, and the ubiquitous NQ320K dataset is built from Natural Questions through diverse and underspecified preprocessing. We introduce ReDSI, the first open-source DSI implementation supporting all three identifier types, together with a parameterizable and well-documented NQ320K construction pipeline. Experimentally, we achieve results that are competitive with or stronger than previous DSI baselines. Moreover, we conduct extensive experiments under model downscaling, covering retrieval effectiveness, parameter efficiency, training methods and decoding strategies, opening novel directions for future research.

生成式检索可复现性索引优化NQ320K

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。