arXiv:2502.09944cs.LGcs.CL2025-02中稿 · Springer Knowledge…被引 4

通过正负样本对比学习,提升神经主题模型的主题质量与稳定性。

Self-Supervised Learning for Neural Topic Models with Variance-Invariance-Covariance Regularization

  • 用对比学习和方差不变性正则化约束主题表示,防止表征坍塌。
  • 在三个数据集上优于基线与当前最优模型,主题更连贯可读。
  • 适合对主题建模、自监督学习感兴趣的学者与工程师。

本文提出一种结合神经主题模型(NTM)与正则化自监督学习的自监督神经主题模型,以提升主题质量。传统主题模型受限于线性假设,而NTM利用神经网络学习文档中隐藏的潜在主题,具备更强灵活性与主题一致性。现有自监督方法采用双分支结构,对同一输入的不同增强版本生成相似嵌入,并通过正则化防止表征坍塌。本文模型通过显式正则化锚点与正样本的潜在主题表示,提升主题表达能力。同时引入对抗性数据增强方法替代启发式采样,并构建基于对比学习的多种变体模型,包含正负样本联合训练机制。在三个数据集上的实验表明,所提模型在定量与定性评估中均优于基线与当前先进模型。

原文摘要 · Abstract (English)

In our study, we propose a self-supervised neural topic model (NTM) that combines the power of NTMs and regularized self-supervised learning methods to improve performance. NTMs use neural networks to learn latent topics hidden behind the words in documents, enabling greater flexibility and the ability to estimate more coherent topics compared to traditional topic models. On the other hand, some self-supervised learning methods use a joint embedding architecture with two identical networks that produce similar representations for two augmented versions of the same input. Regularizations are applied to these representations to prevent collapse, which would otherwise result in the networks outputting constant or redundant representations for all inputs. Our model enhances topic quality by explicitly regularizing latent topic representations of anchor and positive samples. We also introduced an adversarial data augmentation method to replace the heuristic sampling method. We further developed several variation models including those on the basis of an NTM that incorporates contrastive learning with both positive and negative samples. Experimental results on three datasets showed that our models outperformed baselines and state-of-the-art models both quantitatively and qualitatively.

主题建模自监督学习对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。