用大模型+主题建模,自动发现科研文献中的跨领域关联。
Mapping Scientific Literature with Large Language Models and Topic Modeling
- 先用摘要分类主主题,再分析全文找副主题,挖掘隐藏联系。
- 在1500+篇论文上测试,主题多样性高,准确率达75.9%。
- 适合研究趋势分析、跨学科探索的学者使用。
科学文献因学科壁垒、专业术语和稀疏关键词系统而日益碎片化,难以捕捉现代科学的演变结构。本研究提出一种基于大语言模型(LLM)的主题建模框架,用于从主题视角映射科学文献。该方法在20年间的1,500余篇发表于《美国国家科学院院刊》(PNAS)的工程类文章语料上进行验证。采用两阶段分类流程:首先根据摘要为每篇文章分配主要主题类别,随后通过全文分析识别次要类别,揭示语料中潜在的跨主题关联。与传统主题模型不同,该框架生成的主題具有语义可解释性,同时保持优异的量化表现。对比评估显示,其主题多样性更高,主题重叠更低,且在竞争性一致性指标上表现更优。对随机抽样子集的人工验证准确率达75.9%。传统自然语言处理分析进一步确认生成主题与语料中真实语言模式相符。构建的二分图网络连接主次主题,揭示了仅靠摘要或关键词无法察觉的隐含主题关系。结果表明,该框架无需事先了解期刊的双分类体系,即可独立复现其编辑分类结构。总体而言,该方法为科学地图绘制和研究中新兴跨领域联系的发现提供了有力工具。
原文摘要 · Abstract (English)
Scientific literature is increasingly fragmented by disciplinary boundaries, specialized terminology, and potentially sparse keyword systems, making it difficult to capture the evolving structure of modern science. This study introduces a large language model (LLM)-driven framework for mapping scientific literature from a topic modeling perspective. The approach is demonstrated on a 20-year corpus of more than 1,500 engineering-related articles published in the Proceedings of the National Academy of Sciences (PNAS). A two-stage classification pipeline first assigns a primary thematic category to each article based on its abstract, followed by full-text analysis to identify secondary classifications that reveal latent cross-topic connections within the corpus. Unlike conventional topic models, the LLM-based framework produces semantically interpretable topics while maintaining strong quantitative performance. Comparative evaluation against established topic modeling methods shows higher topic diversity and lower overlap with competitive coherence metrics. Manual validation on a randomly sampled subset of abstracts yields an accuracy of 75.9%. Additional traditional natural language processing analyses confirm that the generated topics correspond to meaningful linguistic patterns in the corpus. A bipartite network linking primary and secondary classifications further reveals implicit thematic relationships that are not readily observable through abstracts or keyword systems alone. The findings indicate that the framework independently recovers much of the journal's editorial dual-classification structure without prior knowledge of its schema. Overall, the proposed approach offers a powerful tool for mapping science and identifying emerging cross-topic connections in research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。