arXiv:2409.17383cs.IRcs.AI2024-09被引 27

用语义嵌入与优化搜索提升文档检索准确率

VectorSearch: Enhancing Document Retrieval with Semantic Embeddings and Optimized Search

  • 采用多向量搜索与大模型编码实现语义增强
  • 在真实数据集上显著优于基线方法
  • 适合大规模文档检索场景使用

传统检索方法虽在判断文档相似性方面发挥重要作用,但难以捕捉语义细节。尽管潜在语义分析(LSA)和深度学习有所进展,高维性和语义鸿沟仍导致全面语义理解与精准检索困难。为此,我们提出VectorSearch,结合先进算法、嵌入表示与索引技术,实现更精准的检索。通过创新的多向量搜索操作及基于先进语言模型的查询编码,该方法显著提升了检索准确率。在真实世界数据集上的实验表明,VectorSearch优于基准指标,验证了其在大规模检索任务中的有效性。

原文摘要 · Abstract (English)

Traditional retrieval methods have been essential for assessing document similarity but struggle with capturing semantic nuances. Despite advancements in latent semantic analysis (LSA) and deep learning, achieving comprehensive semantic understanding and accurate retrieval remains challenging due to high dimensionality and semantic gaps. The above challenges call for new techniques to effectively reduce the dimensions and close the semantic gaps. To this end, we propose VectorSearch, which leverages advanced algorithms, embeddings, and indexing techniques for refined retrieval. By utilizing innovative multi-vector search operations and encoding searches with advanced language models, our approach significantly improves retrieval accuracy. Experiments on real-world datasets show that VectorSearch outperforms baseline metrics, demonstrating its efficacy for large-scale retrieval tasks.

文档检索语义嵌入多向量搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。