arXiv:2505.06696cs.CL2025-05被引 1

通过中间层表示提升BERTopic主题建模效果

Enhancing BERTopic with Intermediate Layer Representations

  • 从18种嵌入表示中选择最优中间层特征
  • 在3个数据集上均实现优于默认设置的聚类效果
  • 揭示停用词对不同嵌入配置的影响规律

BERTopic是一种利用基于Transformer的嵌入进行密集聚类的主题建模算法,可有效估计文档语料的主题结构并提取有价值的信息。尽管该方法强大,但嵌入表示的构建方式多样,包括从模型中间层提取表征并施加变换。本研究评估了18种不同的嵌入表示,并在三个多样化数据集上开展实验。通过报告主题连贯性与主题多样性指标,结果表明:在每个数据集上均可找到性能优于BERTopic默认设置的嵌入配置。此外,我们还探究了停用词对不同嵌入配置的影响。

原文摘要 · Abstract (English)

BERTopic is a topic modeling algorithm that leverages transformer-based embeddings to create dense clusters, enabling the estimation of topic structures and the extraction of valuable insights from a corpus of documents. This approach allows users to efficiently process large-scale text data and gain meaningful insights into its structure. While BERTopic is a powerful tool, embedding preparation can vary, including extracting representations from intermediate model layers and applying transformations to these embeddings. In this study, we evaluate 18 different embedding representations and present findings based on experiments conducted on three diverse datasets. To assess the algorithm's performance, we report topic coherence and topic diversity metrics across all experiments. Our results demonstrate that, for each dataset, it is possible to find an embedding configuration that performs better than the default setting of BERTopic. Additionally, we investigate the influence of stop words on different embedding configurations.

主题建模BERT嵌入表示NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。