arXiv:2602.22224cs.IRcs.AI2026-02AAAI被引 2

将半万亿词文本转为高效神经检索系统,单机低延迟运行

DS SERVE: A Framework for Efficient and Scalable Neural Retrieval

论文配图:DS SERVE: A Framework for Efficient and Scalable Neural Retrieval
图 1 · 摘自论文原文
  • 把海量文本压缩成可快速检索的神经索引,支持灵活调整速度与精度
  • 在单机上实现低延迟,内存开销小,可处理半万亿词规模数据
  • 适合大模型检索增强生成、训练数据溯源等场景

我们提出DS-Serve,一个将包含半万亿词的大规模文本数据集转化为高性能神经检索系统的框架。该框架提供网页界面和API接口,在单机上实现低延迟,并具有较小的内存开销。同时支持推理时在延迟、准确率和结果多样性之间的权衡。预计该框架将在大规模检索增强生成(RAG)、训练数据归属分析、训练搜索代理等多种应用场景中发挥重要作用。

原文摘要 · Abstract (English)

We present DS-Serve, a framework that transforms large-scale text datasets, comprising half a trillion tokens, into a high-performance neural retrieval system. DS-Serve offers both a web interface and API endpoints, achieving low latency with modest memory overhead on a single node. The framework also supports inference-time trade-offs between latency, accuracy, and result diversity. We anticipate that DS-Serve will be broadly useful for a range of applications, including large-scale retrieval-augmented generation (RAG), training data attribution, training search agents, and beyond.

神经检索大模型RAG高效系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。