arXiv:2508.13394cs.IR2025-08被引 1

用关键词和概念增强检索,让科学论文搜索更准更快。

CASPER: Concept-integrated Sparse Representation for Scientific Retrieval

  • 以词元和关键词为维度构建稀疏表示,融合细粒度与概念级匹配。
  • 在8个科学检索基准上超越主流稠密与稀疏基线模型。
  • 可解释性强,还能高效生成关键词,适合科研人员快速定位文献。

精准识别相关研究概念对科学检索至关重要。然而,主流稀疏检索方法往往缺乏概念感知的表示能力。为此,我们提出CASPER,一种面向科学搜索的稀疏检索模型,将词元和关键词作为表示单元(即稀疏嵌入空间中的维度),使查询与文档能通过研究概念进行表达,并在细粒度与概念层面实现匹配。此外,我们利用大量学术引用信息(包括标题、引用上下文、作者标注关键词及共引)构建训练数据,捕捉研究概念在不同语境下的表达方式。实验表明,CASPER在8个科学检索基准上均优于强大多密集与稀疏检索基线。我们还通过表示剪枝探索了效果-效率权衡,并验证了其可解释性——证明其可作为高效准确的关键词生成模型。

原文摘要 · Abstract (English)

Identifying relevant research concepts is crucial for effective scientific search. However, primary sparse retrieval methods often lack concept-aware representations. To address this, we propose CASPER, a sparse retrieval model for scientific search that utilizes both tokens and keyphrases as representation units (i.e., dimensions in the sparse embedding space). This enables CASPER to represent queries and documents via research concepts and match them at both granular and conceptual levels. Furthermore, we construct training data by leveraging abundant scholarly references (including titles, citation contexts, author-assigned keyphrases, and co-citations), which capture how research concepts are expressed in diverse settings. Empirically, CASPER outperforms strong dense and sparse retrieval baselines across eight scientific retrieval benchmarks. We also explore the effectiveness-efficiency trade-off via representation pruning and demonstrate CASPER's interpretability by showing that it can serve as an effective and efficient keyphrase generation model.

科学检索稀疏表示关键词生成信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。