arXiv:2506.07747cs.LGcs.CL2025-06ICML被引 1

提出首个可解释且有理论保证的LDA推断算法,速度比现有方法快得多。

E-LDA: Toward Interpretable LDA Topic Models with Strong Guarantees in Logarithmic Parallel Time

  • 采用无梯度组合方法,实现对文档主题的精确推断。
  • 在对数级并行时间内收敛,速度远超现有算法。
  • 结果可解释且支持因果推断,适合社会科学应用。

本文首次为推断LDA主题模型中每个文档的主题分配问题提供了具有可证明保证的实用算法。该问题是社会科学研究、数据探索和因果推断中的核心任务。我们通过提出一种新颖的非梯度、组合式主题建模方法实现这一目标,使算法能在对数级并行计算时间(适应性)内收敛至近似最优后验概率,相比已有方法呈指数级加速。此外,我们的方法能提供可解释性保证,使每个学习到的主题与特定关键词形式关联。最后,与现有方法不同,本方法可保持独立性假设,从而支持将学习到的主题模型用于下游因果推断分析,使研究者能将主题视为干预变量。在实际性能上,我们在多种文本数据集和评估参数下,始终优于当前主流的LDA算法、神经主题模型及基于大语言模型的主题方法,生成的主题语义质量更高。

原文摘要 · Abstract (English)

In this paper, we provide the first practical algorithms with provable guarantees for the problem of inferring the topics assigned to each document in an LDA topic model. This is the primary inference problem for many applications of topic models in social science, data exploration, and causal inference settings. We obtain this result by showing a novel non-gradient-based, combinatorial approach to estimating topic models. This yields algorithms that converge to near-optimal posterior probability in logarithmic parallel computation time (adaptivity) -- exponentially faster than any known LDA algorithm. We also show that our approach can provide interpretability guarantees such that each learned topic is formally associated with a known keyword. Finally, we show that unlike alternatives, our approach can maintain the independence assumptions necessary to use the learned topic model for downstream causal inference methods that allow researchers to study topics as treatments. In terms of practical performance, our approach consistently returns solutions of higher semantic quality than solutions from state-of-the-art LDA algorithms, neural topic models, and LLM-based topic models across a diverse range of text datasets and evaluation parameters.

LDA可解释性因果推断高效算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。