用大模型生成论文多维度摘要,构建更清晰的科学分类体系。
Context-Aware Hierarchical Taxonomy Generation for Scientific Papers via LLM-Guided Multi-Aspect Clustering
- 通过大模型提取论文方法、数据集等多方面特征,分维度生成摘要
- 在11.6千篇论文上构建层次化分类,优于现有方法
- 提供首个专家标注的156个分类基准,适合研究文献组织者
科学文献的快速增长亟需高效的组织与综述方法。现有分类构建方法依赖无监督聚类或直接提示大语言模型(LLMs),常缺乏一致性与细粒度。本文提出一种上下文感知的层级分类生成框架,结合大模型引导的多维度编码与动态聚类。该方法利用大模型识别每篇论文的关键方面(如方法、数据集、评估),生成对应方面的论文摘要,并在各维度上进行编码与聚类,形成连贯的层次结构。此外,我们构建了一个包含156个专家手工制作分类的全新评估基准,涵盖11.6k篇论文,为该任务提供首个自然标注数据集。实验表明,本方法在分类一致性、细粒度和可解释性上均显著优于先前方法,达到当前最佳性能。
原文摘要 · Abstract (English)
The rapid growth of scientific literature demands efficient methods to organize and synthesize research findings. Existing taxonomy construction methods, leveraging unsupervised clustering or direct prompting of large language models (LLMs), often lack coherence and granularity. We propose a novel context-aware hierarchical taxonomy generation framework that integrates LLM-guided multi-aspect encoding with dynamic clustering. Our method leverages LLMs to identify key aspects of each paper (e.g., methodology, dataset, evaluation) and generates aspect-specific paper summaries, which are then encoded and clustered along each aspect to form a coherent hierarchy. In addition, we introduce a new evaluation benchmark of 156 expert-crafted taxonomies encompassing 11.6k papers, providing the first naturally annotated dataset for this task. Experimental results demonstrate that our method significantly outperforms prior approaches, achieving state-of-the-art performance in taxonomy coherence, granularity, and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。