用高阶超图建模文档中词语的共现与重复,提升动态主题模型效果
Dynamic Topic Modeling with a Higher-Order Hypergraphical Representation

- 将文档视为连接共现词的超边,词重复强度以节点权重表示
- 在ICLR数据集上优于传统多项式主题模型,显著捕捉语义演变趋势
- 适合研究科学文献演化、社交媒体话题变迁的学者使用
动态主题模型广泛用于分析科学文献、医疗记录和社交媒体中的趋势演变。传统主题模型通过单一概率向量在多项式单纯形上表示每个主题,并在单一概率机制中隐式耦合词的出现与重复,限制了词间的依赖结构,忽视了具有重叠语义的动态语料中重要的高阶交互。为此,我们提出一种文本的超图表示:每个文档被建模为连接所有共现词的超边,重复强度以节点权重编码。该表示自然分离词的出现与重复,并诱导出一种基于超图的新型多项式分布,其非线性归一化依赖于每篇文档的实际词集。在此似然基础上,我们通过结构化低秩分解与显式的时间正则化构建动态主题模型框架。此外,尽管双线性分解和文档特定的非线性归一化带来固有的非凸性,我们仍建立了局部收敛性保证并推导出非渐近误差界。在合成数据和国际学习表征会议(ICLR)语料上的数值实验表明,该方法在各项指标上均持续优于现有基于多项式的主题模型。
原文摘要 · Abstract (English)
Dynamic topic modeling is widely used to analyze evolving trends in scientific literature, medical records, and social media. Traditional topic models represent each topic through a single probability vector on the multinomial simplex and implicitly couple word occurrence and repetition within one probabilistic mechanism. However, this formulation restricts the dependence structure among words and overlooks informative higher-order interactions, particularly in dynamic corpora with overlapping semantics. To address these limitations, we introduce a hypergraph representation of text where each document is modeled as a hyperedge connecting all co-occurring words, with repetition intensities encoded as node weights. This representation naturally separates word occurrence from repetition and induces a novel hypergraph-based multinomial distribution with a nonlinear normalization depending on the observed word set of each document. Building on this likelihood, we develop a dynamic topic modeling framework via structured low-rank factorizations with explicit temporal regularization on topic-word profiles. Moreover, we establish local convergence guarantees and derive non-asymptotic error bounds despite the intrinsic nonconvexity induced by bilinear factorization and document-specific nonlinear normalization. Numerical experiments on synthetic data and an application to the International Conference on Learning Representations (ICLR) corpus demonstrate consistent improvements over existing multinomial-based topic models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。