arXiv:2608.19133cs.CL2026-08

用嵌入向量分析127亿条评论,发现政治话题语义在持续漂移。

Comment-level Topic Drift Analysis in the Reddit Corpus

论文配图:Comment-level Topic Drift Analysis in the Reddit Corpus
图 1 · 摘自论文原文
  • 用预训练模型生成评论嵌入,通过无监督方法追踪语义演变
  • 政治类话题的语义距离随时间系统性变化,音乐体育则较稳定
  • 提出新检验方法,过滤虚假动态,适合社会计算研究者

我们提出一种基于嵌入的动态主题建模新应用,用于检测和量化大规模语料库中评论级别的主题漂移。通过利用预训练语言模型为短文本生成上下文语义嵌入,我们分析了覆盖2006至2022年、总量达127亿条的Reddit评论。基于这些嵌入,采用无监督方法识别随时间动态演化的主题聚类。主要贡献在于提出在嵌入空间内分析语义漂移与话语演化的框架。同时展示了对现有方法的改进,使其可扩展至大规模数据,并提出并验证了一种零模型对比检验法以过滤虚假动态。关键发现表明,政治与社会敏感话题在嵌入空间中表现出显著的方向性漂移,跨主题距离随时间系统性变化,超出零模型解释范围;而音乐、体育等领域则相对稳定。

原文摘要 · Abstract (English)

We present a novel application of embedding-based dynamic topic modeling techniques to detect and quantify topic drift at the comment level in a massive corpus. By leveraging pretrained language models to generate contextualized semantic embeddings for short text, we analyzed 12.7 billion Reddit comments spanning 2006 to 2022. Using unsupervised methods on these embeddings, we identify dynamically evolving topic clusters over time. Our primary contribution is a methodology for analysis of semantic drift and discourse evolution in the embedding space itself. We also demonstrate modifications to existing methods that enable this analysis at scale, and we propose and demonstrate a null model comparison test to filter spurious dynamics. Key findings suggest that politically and socially contentious topics exhibit significant directional drift in embedding space, with inter-topic distances changing systematically over time beyond what the null model can explain, whereas domains such as music and sports remain comparatively stable.

主题建模语义漂移社交媒体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。