arXiv:2506.22309cs.AIcs.CL2025-06

用形式概念分析提升文本主题的可解释性与可视化

Conceptual Topic Aggregation

  • 基于形式概念分析构建层次化主题结构
  • 在ETYNTKE数据集上比传统方法更易理解数据构成
  • 适合需要深度洞察文本结构的研究者使用

数据爆炸式增长使得人工检查变得不可行,亟需计算方法实现高效数据探索。主题建模已成为分析大规模文本数据的强大工具,能够提取潜在语义结构。然而,现有方法常难以提供可解释的主题表示,限制了对数据结构与内容的深入理解。本文提出FAT-CAT,一种基于形式概念分析(FCA)的方法,用于增强主题聚合与可视化。该方法可处理多样主题与文件类型,按目录分组构建概念格,呈现主题分布的结构化、层次化表示。在ETYNTKE数据集上的案例研究显示,相较其他表示方法,基于FCA的聚合能提供更有意义、更可解释的数据组成洞察。

原文摘要 · Abstract (English)

The vast growth of data has rendered traditional manual inspection infeasible, necessitating the adoption of computational methods for efficient data exploration. Topic modeling has emerged as a powerful tool for analyzing large-scale textual datasets, enabling the extraction of latent semantic structures. However, existing methods for topic modeling often struggle to provide interpretable representations that facilitate deeper insights into data structure and content. In this paper, we propose FAT-CAT, an approach based on Formal Concept Analysis (FCA) to enhance meaningful topic aggregation and visualization of discovered topics. Our approach can handle diverse topics and file types -- grouped by directories -- to construct a concept lattice that offers a structured, hierarchical representation of their topic distribution. In a case study on the ETYNTKE dataset, we evaluate the effectiveness of our approach against other representation methods to demonstrate that FCA-based aggregation provides more meaningful and interpretable insights into dataset composition than existing topic modeling techniques.

主题建模形式概念分析可解释性数据探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。