arXiv:2501.04006cs.IR2025-01

用生成模型提升语义搜索精度,实测相关性达0.905

Advancing Similarity Search with GenAI: A Retrieval Augmented Generation Approach

  • 用生成模型理解上下文,捕捉细微语义信息
  • 在生物医学数据集上相关性达0.905,优于此前方法
  • 温度0.5、20个示例提示时效果最佳,适合检索研究者

本文提出一种新型检索增强生成方法,用于改进语义相似度搜索。该方法利用生成模型捕捉细微语义信息,并基于深度上下文理解生成相似度评分。研究聚焦于包含100对句子的BIOSSES数据集(来自生物医学领域),其相似度搜索相关性结果优于此前水平。通过深入分析模型敏感性,发现最优条件为温度设置0.5、提示中提供20个示例,此时达到最高相似度搜索准确率,具体表现为皮尔逊相关系数高达0.905。研究结果表明生成模型在语义信息检索中具有巨大潜力,为相似度搜索提供了新方向。

原文摘要 · Abstract (English)

This article introduces an innovative Retrieval Augmented Generation approach to similarity search. The proposed method uses a generative model to capture nuanced semantic information and retrieve similarity scores based on advanced context understanding. The study focuses on the BIOSSES dataset containing 100 pairs of sentences extracted from the biomedical domain, and introduces similarity search correlation results that outperform those previously attained on this dataset. Through an in-depth analysis of the model sensitivity, the research identifies optimal conditions leading to the highest similarity search accuracy: the results reveals high Pearson correlation scores, reaching specifically 0.905 at a temperature of 0.5 and a sample size of 20 examples provided in the prompt. The findings underscore the potential of generative models for semantic information retrieval and emphasize a promising research direction to similarity search.

语义搜索生成模型检索增强生物医学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。