用大模型理解科学查询,结合概念索引提升论文检索准确率
Scientific Paper Retrieval with LLM-Guided Semantic-Based Ranking

- 大模型提取查询中的核心科学概念,增强语义匹配
- 在多个基准上优于现有方法,精度显著提升
- 适合需要精准文献发现的研究者使用
科学论文检索对支持文献发现和研究至关重要。尽管密集检索方法在通用任务中表现良好,但难以捕捉对理解科学查询至关重要的细粒度科学概念。近期研究虽利用大语言模型(LLMs)改进查询理解,但往往缺乏与语料库特定知识的结合,可能导致生成内容不可靠或不忠实。为此,我们提出SemRank框架,将大模型引导的查询理解与基于概念的语义索引相结合。每篇论文通过多粒度科学概念(包括通用研究主题和详细关键词)进行索引。查询时,大模型从语料库中识别出核心概念,明确捕捉查询的信息需求。这些概念实现精确语义匹配,显著提升检索准确性。实验表明,SemRank持续改进多种基础检索器的性能,超越强基线方法,且保持高效。
原文摘要 · Abstract (English)
Scientific paper retrieval is essential for supporting literature discovery and research. While dense retrieval methods demonstrate effectiveness in general-purpose tasks, they often fail to capture fine-grained scientific concepts that are essential for accurate understanding of scientific queries. Recent studies also use large language models (LLMs) for query understanding; however, these methods often lack grounding in corpus-specific knowledge and may generate unreliable or unfaithful content. To overcome these limitations, we propose SemRank, an effective and efficient paper retrieval framework that combines LLM-guided query understanding with a concept-based semantic index. Each paper is indexed using multi-granular scientific concepts, including general research topics and detailed key phrases. At query time, an LLM identifies core concepts derived from the corpus to explicitly capture the query's information need. These identified concepts enable precise semantic matching, significantly enhancing retrieval accuracy. Experiments show that SemRank consistently improves the performance of various base retrievers, surpasses strong existing LLM-based baselines, and remains highly efficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。