提出离散马尔可夫桥框架,提升离散数据建模的表达能力。
Discrete Markov Bridge
- 通过矩阵学习与得分学习构建新框架,突破固定转移矩阵限制。
- 在Text8上达到1.38的ELBO,优于现有基线;CIFAR-10表现媲美专用生成模型。
- 理论证明收敛性与空间复杂度可控,适合实际部署。
离散扩散模型近年来在离散数据建模中展现出巨大潜力。然而,现有方法通常在训练中依赖固定的转移矩阵,不仅限制了潜在表示的表达能力——这是变分方法的核心优势之一——也约束了整体设计空间。为解决这些问题,我们提出了离散马尔可夫桥(Discrete Markov Bridge),一种专为离散表示学习设计的新框架。该方法基于两个核心组件:矩阵学习与得分学习。我们进行了严格的理论分析,为矩阵学习建立了正式的性能保证,并证明了整体框架的收敛性。此外,我们分析了方法的空间复杂度,解决了先前研究中识别出的实际约束问题。大量实证评估验证了所提方法的有效性:在Text8数据集上,其证据下界(ELBO)达到1.38,优于已有的基线模型;在CIFAR-10数据集上,性能与图像专用生成方法相当。
原文摘要 · Abstract (English)
Discrete diffusion has recently emerged as a promising paradigm in discrete data modeling. However, existing methods typically rely on a fixed rate transition matrix during training, which not only limits the expressiveness of latent representations, a fundamental strength of variational methods, but also constrains the overall design space. To address these limitations, we propose Discrete Markov Bridge, a novel framework specifically designed for discrete representation learning. Our approach is built upon two key components: Matrix Learning and Score Learning. We conduct a rigorous theoretical analysis, establishing formal performance guarantees for Matrix Learning and proving the convergence of the overall framework. Furthermore, we analyze the space complexity of our method, addressing practical constraints identified in prior studies. Extensive empirical evaluations validate the effectiveness of the proposed Discrete Markov Bridge, which achieves an Evidence Lower Bound (ELBO) of 1.38 on the Text8 dataset, outperforming established baselines. Moreover, the proposed model demonstrates competitive performance on the CIFAR-10 dataset, achieving results comparable to those obtained by image-specific generation approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。