arXiv:2512.18119stat.MEcs.CL2025-12

针对社会科学大数据主题分析难题,提出高效精准的分布式不对称分配模型

Distributed Asymmetric Allocation: A Topic Model for Large Imbalanced Corpora in Social Sciences

  • 采用分布式计算与不对称分配机制加速主题建模
  • 在联合国讲话文本上比LDA快且准确率显著提升
  • 适合需要快速挖掘政治敏感话题的社会科学研究者

社会科学家常使用潜在狄利克雷分配(LDA)在大规模语料中发现特定主题,但面临三大挑战:(1) LDA在大语料上训练耗时长;(2) 无监督LDA在短文档中易将主题碎片化为子主题;(3) 半监督LDA难以识别由种子词定义的具体主题。为此,本文提出一种新型主题模型——分布式不对称分配(DAA),融合多种算法,高效识别大规模语料中重要主题的相关句子。通过在1991至2017年联合国大会演讲文本上应用DAA,结果表明:相比传统LDA,DAA在分类准确性和速度上均有显著提升。研究还揭示,优化LDA的狄利克雷先验对内容分析的准确性至关重要。

原文摘要 · Abstract (English)

Social scientists employ latent Dirichlet allocation (LDA) to find highly specific topics in large corpora, but they often struggle in this task because (1) LDA, in general, takes a significant amount of time to fit on large corpora; (2) unsupervised LDA fragments topics into sub-topics in short documents; (3) semi-supervised LDA fails to identify specific topics defined using seed words. To solve these problems, I have developed a new topic model called distributed asymmetric allocation (DAA) that integrates multiple algorithms for efficiently identifying sentences about important topics in large corpora. I evaluate the ability of DAA to identify politically important topics by fitting it to the transcripts of speeches at the United Nations General Assembly between 1991 and 2017. The results show that DAA can classify sentences significantly more accurately and quickly than LDA thanks to the new algorithms. More generally, the results demonstrate that it is important for social scientists to optimize Dirichlet priors of LDA to perform content analysis accurately.

主题模型社会科学研究分布式计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。