arXiv:2506.12385cs.CLcs.AI2025-06被引 2

用AI从海量论文中挖掘跨领域新发现,加速科学突破。

Recent Advances and Future Directions in Literature-Based Discovery

  • 构建知识图谱+深度学习,自动发现不同领域间的隐藏关联。
  • 引入大模型后,发现效率提升,但仍需大量人工校对。
  • 适合生物医学、药物研发等需要跨学科洞察的研究者。

科学论文的爆炸式增长带来了知识整合与假设生成的迫切需求。文献基础发现(LBD)通过揭示不同领域间此前未知的关联来应对这一挑战。本文综述了2000年以来LBD在三个关键方向上的进展:知识图谱构建、深度学习方法以及预训练模型和大语言模型(LLMs)的融合应用。尽管已有显著进展,但在可扩展性、对结构化数据的依赖及需大量人工标注等方面仍存在根本性挑战。通过分析当前进展并展望未来方向,本文强调了大语言模型在提升LBD能力中的变革作用,旨在帮助研究人员和实践者利用这些技术加速科学创新。

原文摘要 · Abstract (English)

The explosive growth of scientific publications has created an urgent need for automated methods that facilitate knowledge synthesis and hypothesis generation. Literature-based discovery (LBD) addresses this challenge by uncovering previously unknown associations between disparate domains. This article surveys recent methodological advances in LBD, focusing on developments from 2000 to the present. We review progress in three key areas: knowledge graph construction, deep learning approaches, and the integration of pre-trained and large language models (LLMs). While LBD has made notable progress, several fundamental challenges remain unresolved, particularly concerning scalability, reliance on structured data, and the need for extensive manual curation. By examining ongoing advances and outlining promising future directions, this survey underscores the transformative role of LLMs in enhancing LBD and aims to support researchers and practitioners in harnessing these technologies to accelerate scientific innovation.

文献发现大模型知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。