arXiv:2410.18486stat.MEcs.LG2024-10被引 2

用时间泊松分解模型分析30年政坛话语变迁,捕捉话题演变与词汇更替。

Evolving Voices Based on Temporal Poisson Factorisation

  • 基于泊松因子分解扩展出时序模型,处理带时间戳的文本计数数据。
  • 在1981-2016年美国参议院18场演讲中,发现多个话题随时间呈现自回归或随机游走变化。
  • 采用变分推断结合自动微分与文档批处理,支持高效建模高维稀疏文本数据。

世界在变化,讨论话题所用词汇也在演进。分析超过30年的政治演讲数据需要灵活的主题模型,以揭示潜在主题随时间的流行度变化及其词汇演变。本文提出时间泊松因子分解(Temporal Poisson Factorisation, TPF)模型,作为泊松因子分解的扩展,用于建模基于词袋假设和时间戳的稀疏计数数据矩阵。我们探讨并实证比较了不同时间可变潜变量的模型设定,包括一阶自回归结构与随机游走结构。估计基于变分推断,采用坐标上升与文档批处理相结合的自动微分方法。提出合适的变分族以简化推断。对比了对时间可变潜变量使用独立单变量与多变量变分分布的结果。详细分析了在1981–2016年美国参议院18次会期演讲数据上的应用效果。

原文摘要 · Abstract (English)

The world is evolving and so is the vocabulary used to discuss topics in speech. Analysing political speech data from more than 30 years requires the use of flexible topic models to uncover the latent topics and their change in prevalence over time as well as the change in the vocabulary of the topics. We propose the temporal Poisson factorisation (TPF) model as an extension to the Poisson factorisation model to model sparse count data matrices obtained based on the bag-of-words assumption from text documents with time stamps. We discuss and empirically compare different model specifications for the time-varying latent variables consisting either of a flexible auto-regressive structure of order one or a random walk. Estimation is based on variational inference where we consider a combination of coordinate ascent updates with automatic differentiation using batching of documents. Suitable variational families are proposed to ease inference. We compare results obtained using independent univariate variational distributions for the time-varying latent variables to those obtained with a multivariate variant. We discuss in detail the results of the TPF model when analysing speeches from 18 sessions in the U.S. Senate (1981-2016).

主题模型时间序列语音分析泊松因子

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。