自动从引用图生成层次化知识分类树,避免人工偏见。
Taxonomy Tree Generation from Citation Graph
- 基于文本与引用结构递归聚类论文,保证语义与结构一致。
- 用大模型迭代生成每层节点核心概念,保持上下文连贯。
- 联合优化聚类与概念生成,提升分类树质量,适合文献综述者。
从引用图构建知识分类体系对组织科学知识、辅助文献综述和识别新兴研究趋势至关重要。然而,人工构建分类费时费力且易受主观偏见影响,常忽略引用较少但关键的论文。本文提出一种端到端框架HiGTL(Hierarchical Graph Taxonomy Learning),支持人类指令或目标主题引导,实现自动层级化分类生成。首先设计层次化引用图聚类方法,结合论文内容与引用结构递归分组,确保聚类在语义与结构上均合理。其次提出新型节点语义化策略,利用预训练大语言模型(LLM)迭代生成各聚类的核心概念,维持层级间语义一致性。为进一步提升性能,构建联合优化框架,同时微调聚类与概念生成模块,使结构准确性和分类质量协同优化。大量实验表明,HiGTL能有效生成结构清晰、高质量的分类树。
原文摘要 · Abstract (English)
Constructing taxonomies from citation graphs is essential for organizing scientific knowledge, facilitating literature reviews, and identifying emerging research trends. However, manual taxonomy construction is labor-intensive, time-consuming, and prone to human biases, often overlooking pivotal but less-cited papers. In this paper, to enable automatic hierarchical taxonomy generation from citation graphs, we propose HiGTL (Hierarchical Graph Taxonomy Learning), a novel end-to-end framework guided by human-provided instructions or preferred topics. Specifically, we propose a hierarchical citation graph clustering method that recursively groups related papers based on both textual content and citation structure, ensuring semantically meaningful and structurally coherent clusters. Additionally, we develop a novel taxonomy node verbalization strategy that iteratively generates central concepts for each cluster, leveraging a pre-trained large language model (LLM) to maintain semantic consistency across hierarchical levels. To further enhance performance, we design a joint optimization framework that fine-tunes both the clustering and concept generation modules, aligning structural accuracy with the quality of generated taxonomies. Extensive experiments demonstrate that HiGTL effectively produces coherent, high-quality taxonomies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。