构建跨学科论文层级结构,帮助快速把握科研进展与空白领域。
From Papers to Panoramas: Building Hierarchies of Scientific Literature at Scale
- 结合嵌入聚类与大模型提示,实现高效精准的文献分层。
- 在速度与质量平衡上优于纯大模型方法,支持大规模应用。
- 适合科研人员快速导航文献,发现研究热点与未充分探索方向。
科学知识快速增长,使得追踪跨学科进展和高层次概念关联变得困难。尽管引文网络和搜索引擎能检索相关论文,却难以捕捉子领域内研究活动的密度与结构。本文旨在构建覆盖多个抽象层次的高质量科学文献层级结构——从宽泛领域到具体研究。该结构可揭示哪些领域已深入研究、哪些仍待探索。我们提出一种混合方法:结合高效的嵌入聚类与大模型提示,在可扩展性与语义精度间取得平衡。相比依赖大模型迭代构建树结构的方法,本方法在质量-速度权衡上表现更优。所生成的层级结构反映现代科学的跨学科性与多维贡献特征。通过评估大模型代理在层级中定位目标论文的能力,验证其有效性:结果表明该方法显著提升可解释性,并为探索科学文献提供了超越传统搜索的新路径。代码、数据与演示已公开:https://github.com/JHU-CLSP/science-hierarchography
原文摘要 · Abstract (English)
Scientific knowledge is growing rapidly, making it difficult to track progress and high-level conceptual links across broad disciplines. While tools like citation networks and search engines help retrieve related papers, they lack the abstraction needed to capture the density and structure of activity across subfields. We motivate the goal of organizing broad swaths of scientific literature into a high-quality hierarchical structure that spans multiple levels of abstraction---from broad domains to specific studies. Such a representation can provide insights into which fields are well-explored and which are under-explored. To achieve this goal, we develop a hybrid approach that combines efficient embedding-based clustering with LLM-based prompting, striking a balance between scalability and semantic precision. Compared to LLM-heavy methods like iterative tree construction, our approach achieves superior quality-speed trade-offs. Our hierarchies capture different dimensions of research contributions, reflecting the interdisciplinary and multifaceted nature of modern science. We evaluate its utility by measuring how effectively an LLM-based agent can navigate the hierarchy to locate target papers. Results show that our method improves interpretability and offers an alternative pathway for exploring scientific literature beyond traditional search methods. Code, data and demo are available: https://github.com/JHU-CLSP/science-hierarchography
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。