arXiv:2511.07809cs.LGcs.SI2025-11

超大规模政治文本分析新方法,可处理百亿级文档

Analyzing Political Text at Scale with Online Tensor LDA

  • 提出可扩展的张量LDA模型,参数可识别且样本效率高
  • 处理速度超传统LDA 3-4倍,支持百亿文档线性扩展
  • 开源GPU实现,适合研究社交媒体政治议题演化

本文提出一种可线性扩展至百亿文档的主题建模方法。核心贡献包括:1)提出张量隐含狄利克雷分配(TLDA),具备可识别参数与样本复杂度保障;2)计算与内存效率高,运行速度达先前并行化LDA方法的3-4倍以上,可处理超百亿文档数据集;3)提供开源的基于GPU的实现。该方法使以往难以开展的大规模分析成为可能,我们据此完成了两项实际应用研究:首次系统分析了超过两年的推特对话中#MeToo运动的演变过程,以及对2020年总统选举期间选举舞弊话题在社交媒体上的讨论进行深入剖析。该方法为社会科学家提供了近实时分析海量语料库的能力,以回答关键理论问题。

原文摘要 · Abstract (English)

This paper proposes a topic modeling method that scales linearly to billions of documents. We make three core contributions: i) we present a topic modeling method, Tensor Latent Dirichlet Allocation (TLDA), that has identifiable and recoverable parameter guarantees and sample complexity guarantees for large data; ii) we show that this method is computationally and memory efficient (achieving speeds over 3-4x those of prior parallelized Latent Dirichlet Allocation (LDA) methods), and that it scales linearly to text datasets with over a billion documents; iii) we provide an open-source, GPU-based implementation, of this method. This scaling enables previously prohibitive analyses, and we perform two real-world, large-scale new studies of interest to political scientists: we provide the first thorough analysis of the evolution of the #MeToo movement through the lens of over two years of Twitter conversation and a detailed study of social media conversations about election fraud in the 2020 presidential election. Thus this method provides social scientists with the ability to study very large corpora at scale and to answer important theoretically-relevant questions about salient issues in near real-time.

主题建模大规模分析政治文本张量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。