用实体-方面对捕捉科学概念多维度语义,提升文献检索精度
PairSem: LLM-Guided Pairwise Semantic Matching for Scientific Document Retrieval
- 将科学概念建模为实体与方面组成的成对结构,捕捉复杂语义
- 在多个数据集上显著提升检索性能,优于现有密集检索方法
- 无需标注数据,兼容各类检索器,适合科研知识发现场景
科学文献检索是支持跨领域研究和知识发现的关键任务。然而,现有密集检索方法因依赖整体嵌入且领域理解有限,难以捕捉文本中的细粒度科学概念。近期方法利用大语言模型提取细粒度语义实体以增强匹配,但通常将实体视为独立片段,忽略了科学概念的多面性。为此,我们提出成对语义匹配(PairSem)框架,将相关语义表示为实体-方面对,以捕捉复杂、多方面的科学概念。PairSem无需监督、与基础检索器无关且可即插即用,无需查询-文档标签或实体标注即可实现精准、上下文感知的匹配。在多个数据集和检索器上的大量实验表明,PairSem显著提升了检索性能,凸显了在科学信息检索中建模多方面语义的重要性。
原文摘要 · Abstract (English)
Scientific document retrieval is a critical task for enabling knowledge discovery and supporting research across diverse domains. However, existing dense retrieval methods often struggle to capture fine-grained scientific concepts in texts due to their reliance on holistic embeddings and limited domain understanding. Recent approaches leverage large language models (LLMs) to extract fine-grained semantic entities and enhance semantic matching, but they typically treat entities as independent fragments, overlooking the multi-faceted nature of scientific concepts. To address this limitation, we propose Pairwise Semantic Matching (PairSem), a framework that represents relevant semantics as entity-aspect pairs, capturing complex, multi-faceted scientific concepts. PairSem is unsupervised, base retriever-agnostic, and plug-and-play, enabling precise and context-aware matching without requiring query-document labels or entity annotations. Extensive experiments on multiple datasets and retrievers demonstrate that PairSem significantly improves retrieval performance, highlighting the importance of modeling multi-aspect semantics in scientific information retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。