让大模型动态构建多维度学术分类体系,更贴合科研领域演进。
TaxoAdapt: Aligning LLM-Based Multidimensional Taxonomy Construction to Evolving Research Corpora
- 基于大模型生成初始分类,通过迭代聚类动态扩展维度与层级。
- 在多个计算机科学会议数据上,分类粒度提升26.51%,一致性提高50.41%。
- 适合需要追踪科研演变、构建多维知识体系的研究者使用。
科学领域的快速演进给文献组织与检索带来挑战。传统专家构建的分类体系耗时且成本高;现有自动方法或过度依赖特定语料导致泛化性差,或过度依赖大模型预训练知识,忽视科学领域的动态变化。此外,这些方法未能充分考虑科研文献的多维属性——一篇论文可能涉及多种维度(如方法、新任务、评估指标、基准)。为此,我们提出 TaxoAdapt 框架,能动态将大模型生成的分类体系适配到目标语料,在多个维度上进行迭代式层次分类,根据语料主题分布扩展分类宽度与深度。我们在多个年份的计算机科学会议数据上验证其性能,结果表明,TaxoAdapt 构建的分类体系相比最先进基线,粒度保留率提升26.51%,一致性提升50.41%(由大模型评估)。
原文摘要 · Abstract (English)
The rapid evolution of scientific fields introduces challenges in organizing and retrieving scientific literature. While expert-curated taxonomies have traditionally addressed this need, the process is time-consuming and expensive. Furthermore, recent automatic taxonomy construction methods either (1) over-rely on a specific corpus, sacrificing generalizability, or (2) depend heavily on the general knowledge of large language models (LLMs) contained within their pre-training datasets, often overlooking the dynamic nature of evolving scientific domains. Additionally, these approaches fail to account for the multi-faceted nature of scientific literature, where a single research paper may contribute to multiple dimensions (e.g., methodology, new tasks, evaluation metrics, benchmarks). To address these gaps, we propose TaxoAdapt, a framework that dynamically adapts an LLM-generated taxonomy to a given corpus across multiple dimensions. TaxoAdapt performs iterative hierarchical classification, expanding both the taxonomy width and depth based on corpus' topical distribution. We demonstrate its state-of-the-art performance across a diverse set of computer science conferences over the years to showcase its ability to structure and capture the evolution of scientific fields. As a multidimensional method, TaxoAdapt generates taxonomies that are 26.51% more granularity-preserving and 50.41% more coherent than the most competitive baselines judged by LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。