arXiv:2512.23471cs.CLcs.AI2025-12

用大模型嵌入与密度树发现文本的多尺度语义结构

Discovering Multi-Scale Semantic Structure in Text Corpora Using Density-Based Trees and LLM Embeddings

  • 基于嵌入空间密度逐步放松约束,生成层次化语义树
  • 中间密度层级语义对齐最优,突变点对应主题分辨率变化
  • 适合分析大规模动态文本库,揭示跨领域关联与新兴主题

大语言模型使文档可表示为稠密语义嵌入,支持大规模文本集合的相似性操作。然而,多数网络级系统仍依赖扁平聚类或预定义分类体系,限制了对层级主题关系的洞察。本文首次将层次化密度建模应用于大语言模型嵌入,不强制固定分类或单一聚类粒度,而是逐步放松局部密度约束,揭示紧凑语义群如何合并为更广泛的主题区域。生成的树结构直接从数据中编码多尺度语义组织,使主题间的结构性关系显式可见。在标准文本基准上评估显示,语义对齐在中间密度层级达到峰值,而突变点对应有意义的语义分辨率变化。此外,该方法应用于大型机构与科学文献库,揭示主导领域、跨学科邻近性及新兴主题集群。通过将层次结构视为嵌入空间密度的涌现属性,该方法提供了一种可解释、多尺度的语义结构表征,适用于大规模、持续演化的文本集合。

原文摘要 · Abstract (English)

Recent advances in large language models enable documents to be represented as dense semantic embeddings, supporting similarity-based operations over large text collections. However, many web-scale systems still rely on flat clustering or predefined taxonomies, limiting insight into hierarchical topic relationships. In this paper we operationalize hierarchical density modeling on large language model embeddings in a way not previously explored. Instead of enforcing a fixed taxonomy or single clustering resolution, the method progressively relaxes local density constraints, revealing how compact semantic groups merge into broader thematic regions. The resulting tree encodes multi-scale semantic organization directly from data, making structural relationships between topics explicit. We evaluate the hierarchies on standard text benchmarks, showing that semantic alignment peaks at intermediate density levels and that abrupt transitions correspond to meaningful changes in semantic resolution. Beyond benchmarks, the approach is applied to large institutional and scientific corpora, exposing dominant fields, cross-disciplinary proximities, and emerging thematic clusters. By framing hierarchical structure as an emergent property of density in embedding spaces, this method provides an interpretable, multi-scale representation of semantic structure suitable for large, evolving text collections.

语义结构层次聚类嵌入分析多尺度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。