arXiv:2602.22213cs.IRcs.AI2026-02被引 3

用大模型扩充分类体系,解决知识分类覆盖不足问题

Enriching Taxonomies Using Large Language Models

  • 以现有分类体系为种子,用大模型生成候选节点
  • 通过验证机制过滤幻觉,确保新增节点语义相关
  • 支持溯源追踪与可视化,适合知识管理场景

分类体系在各领域信息结构化中至关重要。然而,许多现有分类体系存在覆盖范围有限、节点过时或模糊等问题,影响知识检索效果。为此,我们提出 Taxoria,一种基于大语言模型(LLM)的分类体系增强管道。不同于从大模型中提取内部分类的方法,Taxoria以已有分类体系为种子,通过提示工程引导大模型生成候选节点。这些候选节点经验证后,可有效缓解幻觉并保证语义相关性,最终整合进原分类体系。输出包含带溯源追踪的增强版分类体系,并提供可视化工具用于分析。

原文摘要 · Abstract (English)

Taxonomies play a vital role in structuring and categorizing information across domains. However, many existing taxonomies suffer from limited coverage and outdated or ambiguous nodes, reducing their effectiveness in knowledge retrieval. To address this, we present Taxoria, a novel taxonomy enrichment pipeline that leverages Large Language Models (LLMs) to enhance a given taxonomy. Unlike approaches that extract internal LLM taxonomies, Taxoria uses an existing taxonomy as a seed and prompts an LLM to propose candidate nodes for enrichment. These candidates are then validated to mitigate hallucinations and ensure semantic relevance before integration. The final output includes an enriched taxonomy with provenance tracking and visualization of the final merged taxonomy for analysis.

分类体系大模型知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。