动态分析新闻流中的主题演化,自动识别随时间变化的主题数量。
Stick-Breaking Embedded Topic Model with Continuous Optimal Transport for Online Analysis of Document Streams
- 用截断棒折构造动态主题分布,自动确定每阶段活跃主题数。
- 基于连续最优传输的融合策略,有效合并不同时段的主题嵌入。
- 适用于实时新闻、社交媒体等持续演化的文本流分析。
在线主题模型是用于在持续演化的数据流中识别潜在主题的无监督算法。尽管这类方法更贴近现实场景,但因面临额外挑战,研究关注度远低于离线模型。为此,我们提出SB-SETM,将嵌入式主题模型(ETM)扩展至处理数据流,通过合并连续部分文档批次生成的模型实现。该方法(i)采用截断棒折构造主题-文档分布,使模型能从数据中自动推断各时间步的活跃主题数量;(ii)引入基于连续最优传输的融合策略,适配高维主题空间的特性。数值实验表明,在模拟场景中SB-SETM优于基线方法。我们在涵盖2022–2023年俄乌战争的新闻文章真实语料上进行了广泛测试,验证了其有效性。
原文摘要 · Abstract (English)
Online topic models are unsupervised algorithms to identify latent topics in data streams that continuously evolve over time. Although these methods naturally align with real-world scenarios, they have received considerably less attention from the community compared to their offline counterparts, due to specific additional challenges. To tackle these issues, we present SB-SETM, an innovative model extending the Embedded Topic Model (ETM) to process data streams by merging models formed on successive partial document batches. To this end, SB-SETM (i) leverages a truncated stick-breaking construction for the topic-per-document distribution, enabling the model to automatically infer from the data the appropriate number of active topics at each timestep; and (ii) introduces a merging strategy for topic embeddings based on a continuous formulation of optimal transport adapted to the high dimensionality of the latent topic space. Numerical experiments show SB-SETM outperforming baselines on simulated scenarios. We extensively test it on a real-world corpus of news articles covering the Russian-Ukrainian war throughout 2022-2023.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。