arXiv:2504.10613cs.IR2025-04

用神经检索模型帮生物医学知识库补全多实体关系,提升文献筛选效率。

Enhancing Document Retrieval for Curating N-ary Relations in Knowledge Bases

  • 基于知识库构建弱监督数据集,用对比学习处理不完整关系。
  • 在两个生物医学检索基准上NDCG@10提升5.7和3.7个百分点。
  • 适合从事知识库构建与生物医学信息抽取的研究者。

生物医学知识库的维护依赖从文献中提取准确的多实体关系事实,这一过程仍主要依赖人工和专家判断。其中关键步骤是检索能支持或补全部分已知多实体关系的文献。本文提出一种神经检索模型,用于辅助知识库构建,识别可填补缺失关系参数并提供上下文证据的文档。为减少对稀缺标注数据的依赖,利用现有知识库记录构建弱监督训练集。方法引入两项关键技术:(i) 分层对比损失,使模型能从噪声和不完整的结构中学习;(ii) 平衡采样策略,从多样化知识库记录中生成高质量负样本。在两个生物医学检索基准上,本方法表现达到当前最优,相比强基线在NDCG@10上分别提升5.7和3.7个百分点。

原文摘要 · Abstract (English)

Curation of biomedical knowledge bases (KBs) relies on extracting accurate multi-entity relational facts from the literature - a process that remains largely manual and expert-driven. An essential step in this workflow is retrieving documents that can support or complete partially observed n-ary relations. We present a neural retrieval model designed to assist KB curation by identifying documents that help fill in missing relation arguments and provide relevant contextual evidence. To reduce dependence on scarce gold-standard training data, we exploit existing KB records to construct weakly supervised training sets. Our approach introduces two key technical contributions: (i) a layered contrastive loss that enables learning from noisy and incomplete relational structures, and (ii) a balanced sampling strategy that generates high-quality negatives from diverse KB records. On two biomedical retrieval benchmarks, our approach achieves state-of-the-art performance, outperforming strong baselines in NDCG@10 by 5.7 and 3.7 percentage points, respectively.

知识库文献检索弱监督多实体关系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。