arXiv:2412.06649cs.IRcs.AI2024-12被引 1

用词向量与近似搜索加速海量数据的语义检索

Semantic Search and Recommendation Algorithm

  • 结合Word2Vec与Annoy索引实现语义匹配
  • 100GB数据下仍保持高精度与高性能
  • 适合大规模信息检索场景的系统优化

本文提出一种新型语义搜索算法,采用Word2Vec与Annoy索引技术,提升从大规模数据集中进行信息检索的效率。该方法克服传统搜索方式在速度、准确率和可扩展性方面的局限。在最大达100GB的数据集上测试表明,该方法能够在处理海量数据的同时维持高精度与良好性能。

原文摘要 · Abstract (English)

This paper introduces a new semantic search algorithm that uses Word2Vec and Annoy Index to improve the efficiency of information retrieval from large datasets. The proposed approach addresses the limitations of traditional search methods by offering enhanced speed, accuracy, and scalability. Testing on datasets up to 100GB demonstrates the method's effectiveness in processing vast amounts of data while maintaining high precision and performance.

语义搜索推荐算法向量检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。