负采样让神经主题模型更准,提升主题质量与多样性。
Evaluating Negative Sampling Approaches for Neural Topic Models
- 在变分自编码器解码器中引入负采样,增强正负样本对比学习。
- 在4个数据集上显著提升主题一致性、多样性和文档分类准确率。
- 适合关注主题建模性能优化的研究者和实践者。
负采样作为一种有效技术,通过引入‘学比较’范式,使深度学习模型能更好地学习表征。尽管其在计算机视觉和自然语言处理中已有广泛应用,但在无监督领域如主题建模中的影响尚未深入研究。本文对多种负采样策略在神经主题模型中的效果进行了全面分析。通过在基于变分自编码器的神经主题模型解码器中引入负采样,我们在4个公开数据集上进行实验,结果表明该方法显著提升了主题一致性、主题多样性及文档分类准确率。人工评估也显示,加入负采样后生成的主题质量更高。这些发现凸显了负采样作为提升神经主题模型效能的重要工具的潜力。
原文摘要 · Abstract (English)
Negative sampling has emerged as an effective technique that enables deep learning models to learn better representations by introducing the paradigm of learn-to-compare. The goal of this approach is to add robustness to deep learning models to learn better representation by comparing the positive samples against the negative ones. Despite its numerous demonstrations in various areas of computer vision and natural language processing, a comprehensive study of the effect of negative sampling in an unsupervised domain like topic modeling has not been well explored. In this paper, we present a comprehensive analysis of the impact of different negative sampling strategies on neural topic models. We compare the performance of several popular neural topic models by incorporating a negative sampling technique in the decoder of variational autoencoder-based neural topic models. Experiments on four publicly available datasets demonstrate that integrating negative sampling into topic models results in significant enhancements across multiple aspects, including improved topic coherence, richer topic diversity, and more accurate document classification. Manual evaluations also indicate that the inclusion of negative sampling into neural topic models enhances the quality of the generated topics. These findings highlight the potential of negative sampling as a valuable tool for advancing the effectiveness of neural topic models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。