分离文本中的通用主题与环境特异性词汇,提升跨场景预测与因果推断能力。
Multi-environment Topic Models
- 将主题建模中的全局主题与环境特异性词分开学习
- 在政治文本数据上实现跨环境更优的预测性能,且在分布外数据上表现更强
- 适用于需跨场景分析或因果效应识别的研究者
概率主题模型是挖掘大规模文本数据潜在主题的强大工具。在许多文本数据中,文档还带有协变量(如来源、风格、政治倾向),这些可视为调节“全局”(环境无关)主题表示的环境。准确学习此类表示对在未见环境中预测新文档以及估计主题对现实结果的因果效应至关重要。为此,我们提出多环境主题模型(MTM),一种无监督概率模型,能分离全局与环境特异性词汇。在从广告到推文和演讲的多种政治内容数据上实验表明,MTM生成可解释的全局主题,并呈现鲜明的环境特异性词汇。在多环境数据上,MTM在分布内与分布外均优于强基线,并能发现准确的因果效应。
原文摘要 · Abstract (English)
Probabilistic topic models are a powerful tool for extracting latent themes from large text datasets. In many text datasets, we also observe per-document covariates (e.g., source, style, political affiliation) that act as environments that modulate a "global" (environment-agnostic) topic representation. Accurately learning these representations is important for prediction on new documents in unseen environments and for estimating the causal effect of topics on real-world outcomes. To this end, we introduce the Multi-environment Topic Model (MTM), an unsupervised probabilistic model that separates global and environment-specific terms. Through experimentation on various political content, from ads to tweets and speeches, we show that the MTM produces interpretable global topics with distinct environment-specific words. On multi-environment data, the MTM outperforms strong baselines in and out-of-distribution. It also enables the discovery of accurate causal effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。