通过自适应降维聚类,让短文本生成更易懂的主题。
TopiCLEAR: Adaptive embedding clustering for interpretable topic discovery from short texts
- 用自适应降维识别嵌入空间中的低维结构,聚类生成主题。
- 在4个数据集上与人工标注高度一致,尤其擅长短文本。
- 适合关注可解释性主题发现的研究者和实际应用者。
主题发现是文本挖掘的核心技术,旨在从大规模文档集合中识别抽象主题。近期方法通过聚类预训练语言模型生成的文档或句子嵌入,并将每个聚类表示为一个主题。尽管表现优异,但嵌入空间几何结构与人类可解释主题组织之间的设计原则仍不清晰。本文提出TopiCLEAR(基于自适应降维的嵌入聚类主题发现),融合文档嵌入与迭代聚类,假设人类可解释主题对应嵌入空间中的低维几何结构,并利用自适应降维识别这些结构。我们采用文档级评估方法,结合基于人工标注的数据定量评估与无真实标签时的定性分析。在涵盖正式与非正式文本的四个基准数据集上,TopiCLEAR始终与人工标注高度一致,尤其在短文本上表现突出。对Twitter数据的案例研究显示,其生成的主题比LDA更具可解释性,能恢复人工标注的主题结构及连贯子主题结构。结果表明,在低维主题空间中聚类文档对主题发现有效。
原文摘要 · Abstract (English)
Topic discovery is a fundamental technique for text mining that identifies abstract topics within large document collections. A recent approach to topic discovery is to cluster document or sentence embeddings, typically obtained from pre-trained language models, and represent each cluster as a topic. Despite their strong empirical performance, the design principles linking the geometry of embedding spaces to human-interpretable topic organization remain unclear. Clarifying the relationship between these geometric structures and topic interpretability is therefore a key challenge in topic discovery. In this study, we propose TopiCLEAR (Topic discovery by CLustering Embeddings with Adaptive dimensionality Reduction), a simple framework that integrates document embeddings with iterative clustering based on adaptive dimensionality reduction. TopiCLEAR is guided by the hypothesis that human-interpretable topics correspond to low-dimensional geometric structures in embedding spaces and leverages adaptive dimensionality reduction to identify them. We evaluate topic quality using a document-level approach that combines quantitative evaluation based on human-labeled data with qualitative assessment in the absence of ground-truth labels. Experiments on four benchmark datasets, covering both formal and informal texts, show that TopiCLEAR consistently achieves strong agreement with human annotations, particularly for short and informal texts. Furthermore, a case study on Twitter data demonstrates that TopiCLEAR produces more interpretable topics than Latent Dirichlet Allocation (LDA), recovering both human-annotated topic structure and coherent sub-topic structure. These results highlight the effectiveness of clustering documents in a low-dimensional topic space for topic discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。