通过聚合短文本提升政治文本主题模型可解释性
The study of short texts in digital politics: Document aggregation for topic modeling
- 按账号聚合推文构建更长文档进行建模
- 聚合后主题与各州关联性显著增强
- 适用于研究社交媒体中的政治话语结构
统计主题建模广泛应用于政治学文本分析,研究者常处理从推文到演讲等不同长度的文本。关于文档长度如何影响主题模型可解释性仍存争议。本文研究基于自然划分单位将短文档聚合为更长文档的影响。以2016年4月至2020年9月期间美国各州立法者的100万条推文为例,发现按账号聚合后,主题与具体州的关联性高于使用原始推文;该结果在按出生城市聚合维基百科页面时也得到复现,表明文档定义方式对主题建模结果有显著影响。
原文摘要 · Abstract (English)
Statistical topic modeling is widely used in political science to study text. Researchers examine documents of varying lengths, from tweets to speeches. There is ongoing debate on how document length affects the interpretability of topic models. We investigate the effects of aggregating short documents into larger ones based on natural units that partition the corpus. In our study, we analyze one million tweets by U.S. state legislators from April 2016 to September 2020. We find that for documents aggregated at the account level, topics are more associated with individual states than when using individual tweets. This finding is replicated with Wikipedia pages aggregated by birth cities, showing how document definitions can impact topic modeling results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。