arXiv:2606.26963cs.CL2026-06

从海量异构文本中自动构建可解释的术语层级结构。

Term-Centric Hierarchy Induction from Heterogeneous Corpora

论文配图:Term-Centric Hierarchy Induction from Heterogeneous Corpora
图 1 · 摘自论文原文
  • 以术语为中心,通过自动提取实现跨源文档对齐。
  • 在百万级多源数据上提升跨源一致性和层级质量。
  • 适合政策分析、技术地图绘制等实际应用。

从多样文本来源中组织知识并构建可解释的层级结构,对于政策分析、创新监测和探索性领域映射等任务至关重要。现有分类法归纳方法通常依赖文档级表示,捕捉的是整篇文档内容而非知识组织所需的特定领域概念,限制了其在异构来源间的泛化能力。本文提出一种以术语为中心的框架,可在大规模文档集合上诱导层次化分类体系。该方法利用自动术语提取将不同来源的文档映射到共享表示空间,实现鲁棒的跨源对齐;在此基础上,结合领域先验与数据驱动聚类构建可解释的层级结构。在超过一百万份英文和德文多源文档的新基准上进行实验,结果表明该方法在跨源一致性与层级质量方面均优于基于文本和摘要的基线模型。针对德国区域创新分析的案例研究进一步验证了其在技术图谱绘制中的实用价值。

原文摘要 · Abstract (English)

Organizing knowledge from diverse text sources into interpretable hierarchies is crucial for tasks such as policy analysis, innovation monitoring, and exploratory domain mapping. Existing taxonomy induction methods typically rely on document-level representations that capture entire documents rather than the specific domain concepts relevant for knowledge organization, limiting their ability to generalize across heterogeneous sources. We propose a term-centric framework for inducing hierarchical taxonomies from heterogeneous corpora that scales to massive document collections. Our approach maps documents from diverse sources into a shared representation space using automatic term extraction, enabling robust cross-source alignment. Based on these representations, we construct interpretable hierarchies that integrate domain priors with datadriven clustering. Experiments on a novel English and German multi-source benchmark of over one million documents demonstrate that our method improves cross-source coherence and hierarchy quality over text- and summarybased baselines. A case study on German regional innovation analysis further demonstrates its practical utility for technology landscape mapping.

知识组织术语挖掘层级结构多源分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。