用AI整合论文数据与关系,让科研探索更智能
Intelligent Scientific Literature Explorer using Machine Learning (ISLE)
- 融合关键词与语义搜索,提升文献召回准确率
- 构建包含作者、机构等的多层知识图谱,揭示研究关联
- 适合需要快速掌握领域脉络的研究者使用
科学出版物的快速增长给研究人员发现、理解与解读相关文献带来挑战。传统关键词搜索缺乏语义理解,现有AI工具多聚焦于单一任务如检索或聚类。本文提出一个集成式科学文献探索系统,整合大规模数据获取、混合检索、语义主题建模和异构知识图谱构建。系统通过合并arXiv全文数据与OpenAlex结构化元数据构建全面语料库。采用基于BM25与嵌入向量的混合检索架构,使用倒数排名融合(Reciprocal Rank Fusion)实现融合。根据计算资源选择BERTopic或非负矩阵分解进行主题建模。知识图谱将论文、作者、机构、国家及提取的主题统一为可解释结构。系统提供多层级探索环境,不仅返回相关文献,还揭示查询背后的概念与关系网络。在多个查询上的评估显示,该系统在检索相关性、主题一致性与可解释性方面均有提升。该框架为人工智能辅助科学发现提供了可扩展的基础。
原文摘要 · Abstract (English)
The rapid acceleration of scientific publishing has created substantial challenges for researchers attempting to discover, contextualize, and interpret relevant literature. Traditional keyword-based search systems provide limited semantic understanding, while existing AI-driven tools typically focus on isolated tasks such as retrieval, clustering, or bibliometric visualization. This paper presents an integrated system for scientific literature exploration that combines large-scale data acquisition, hybrid retrieval, semantic topic modeling, and heterogeneous knowledge graph construction. The system builds a comprehensive corpus by merging full-text data from arXiv with structured metadata from OpenAlex. A hybrid retrieval architecture fuses BM25 lexical search with embedding-based semantic search using Reciprocal Rank Fusion. Topic modeling is performed on retrieved results using BERTopic or non-negative matrix factorization depending on computational resources. A knowledge graph unifies papers, authors, institutions, countries, and extracted topics into an interpretable structure. The system provides a multi-layered exploration environment that reveals not only relevant publications but also the conceptual and relational landscape surrounding a query. Evaluation across multiple queries demonstrates improvements in retrieval relevance, topic coherence, and interpretability. The proposed framework contributes an extensible foundation for AI-assisted scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。