arXiv:2507.03329cs.AI2025-07

专为神经科学设计的向量模型,提升文献检索精度。

NDAI-NeuroMAP: A Neuroscience-Specific Embedding Model for Domain-Specific Retrieval

  • 构建50万组神经科学三元组训练数据,融合定义与知识图谱。
  • 在2.4万条查询上超越现有通用模型,召回率显著提升。
  • 适合临床NLP和神经科学知识库检索场景使用。

我们提出NDAI-NeuroMAP,首个专为神经科学领域设计的密集向量嵌入模型,用于高精度信息检索。方法包括构建包含50万组精心构造三元组(查询-正例-负例)的领域专用训练语料库,补充25万条神经科学术语定义及25万条来自权威神经学本体的知识图谱三元组。采用FremyCompany/BioLORD-2023基础模型进行细调,结合对比学习与基于三元组的度量学习,实现多目标优化。在约2.4万条神经科学专属查询组成的保留测试集上评估,性能显著优于当前最优的通用与生物医学嵌入模型。结果表明,针对神经科学的专用嵌入架构对神经科学导向的RAG系统及相关临床自然语言处理应用至关重要。

原文摘要 · Abstract (English)

We present NDAI-NeuroMAP, the first neuroscience-domain-specific dense vector embedding model engineered for high-precision information retrieval tasks. Our methodology encompasses the curation of an extensive domain-specific training corpus comprising 500,000 carefully constructed triplets (query-positive-negative configurations), augmented with 250,000 neuroscience-specific definitional entries and 250,000 structured knowledge-graph triplets derived from authoritative neurological ontologies. We employ a sophisticated fine-tuning approach utilizing the FremyCompany/BioLORD-2023 foundation model, implementing a multi-objective optimization framework combining contrastive learning with triplet-based metric learning paradigms. Comprehensive evaluation on a held-out test dataset comprising approximately 24,000 neuroscience-specific queries demonstrates substantial performance improvements over state-of-the-art general-purpose and biomedical embedding models. These empirical findings underscore the critical importance of domain-specific embedding architectures for neuroscience-oriented RAG systems and related clinical natural language processing applications.

神经科学嵌入模型信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。