通过对比学习提升神经主题模型的主题可解释性
Enhancing Topic Interpretability for Neural Topic Modeling through Topic-wise Contrastive Learning
- 采用主题级对比学习,动态评估并优化主题内部一致性与外部差异性
- 在三个数据集上均显著提升主题可解释性,优于当前主流神经主题模型
- 适合关注主题模型可解释性的研究人员和实际应用者
数据挖掘与知识发现是从海量数据中提取有价值信息的关键环节。神经主题模型(NTMs)作为该领域的重要无监督工具应运而生。然而,主流NTM以最大化数据似然为核心目标,与数据挖掘和知识发现的核心目标——从大规模数据中揭示可解释的洞察——存在偏差。过度强调似然最大化而缺乏主题正则化,会导致主题空间过于宽泛。本文提出一种创新方法,引入对比学习机制来衡量主题可解释性。我们设计了名为ContraTopic的新框架,集成一个可微分正则项,可在训练过程中评估主题的多重可解释性特征。该正则项采用独特的主题级对比策略,促进主题内部一致性及主题间的清晰区分。在三个不同数据集上的全面实验表明,本方法生成的主题在可解释性方面显著优于现有先进方法。
原文摘要 · Abstract (English)
Data mining and knowledge discovery are essential aspects of extracting valuable insights from vast datasets. Neural topic models (NTMs) have emerged as a valuable unsupervised tool in this field. However, the predominant objective in NTMs, which aims to discover topics maximizing data likelihood, often lacks alignment with the central goals of data mining and knowledge discovery which is to reveal interpretable insights from large data repositories. Overemphasizing likelihood maximization without incorporating topic regularization can lead to an overly expansive latent space for topic modeling. In this paper, we present an innovative approach to NTMs that addresses this misalignment by introducing contrastive learning measures to assess topic interpretability. We propose a novel NTM framework, named ContraTopic, that integrates a differentiable regularizer capable of evaluating multiple facets of topic interpretability throughout the training process. Our regularizer adopts a unique topic-wise contrastive methodology, fostering both internal coherence within topics and clear external distinctions among them. Comprehensive experiments conducted on three diverse datasets demonstrate that our approach consistently produces topics with superior interpretability compared to state-of-the-art NTMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。