用图结构提升文档主题建模,让相似文档共享主题分布。
Graph Topic Modeling for Documents with Spatial or Covariate Dependencies
- 将文档视为节点,通过图结构建模文档相似性,约束主题分布。
- 提出快速图正则化SVD算法,误差有概率保证,推理速度更快。
- 适合处理带空间或协变量依赖的文档数据,如地理文本、社交网络内容。
我们研究如何将文档级元数据融入主题建模以提升主题混合估计。针对现有贝叶斯方法计算复杂度高且缺乏理论保障的问题,将频率学派的主题建模框架pLSI扩展为支持文档级协变量或已知文档相似性的图形式。将文档建模为节点,边表示相似性,提出一种基于快速图正则化迭代SVD的新估计器,促使相似文档共享相似的主题混合比例。我们推导了该方法的高概率估计误差界,并设计专用交叉验证方法优化正则化参数。在合成数据集和三个真实语料库上的实验表明,该方法性能更优、推理更快,优于现有贝叶斯方法。
原文摘要 · Abstract (English)
We address the challenge of incorporating document-level metadata into topic modeling to improve topic mixture estimation. To overcome the computational complexity and lack of theoretical guarantees in existing Bayesian methods, we extend probabilistic latent semantic indexing (pLSI), a frequentist framework for topic modeling, by incorporating document-level covariates or known similarities between documents through a graph formalism. Modeling documents as nodes and edges denoting similarities, we propose a new estimator based on a fast graph-regularized iterative singular value decomposition (SVD) that encourages similar documents to share similar topic mixture proportions. We characterize the estimation error of our proposed method by deriving high-probability bounds and develop a specialized cross-validation method to optimize our regularization parameters. We validate our model through comprehensive experiments on synthetic datasets and three real-world corpora, demonstrating improved performance and faster inference compared to existing Bayesian methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。