解决数据缺失时的得分匹配问题,提升小样本低维与高维复杂场景下的建模效果。
Score Matching With Missing Data
- 提出两种新方法:重要性加权与变分推断,适配任意坐标的不完整数据
- 小样本低维场景下重要性加权法表现优异,高维复杂任务中变分方法更优
- 适用于扩散模型、图模型估计等需处理缺失数据的领域
得分匹配是学习数据分布的重要工具,广泛应用于扩散过程、能量模型和图模型估计等领域。然而,现有研究极少关注数据不完整情况下的应用。本文针对部分坐标缺失的灵活场景,拓展了得分匹配及其主要变体。提出两种通用方法:重要性加权(IW)与变分方法。在有限域设定下,为IW方法提供了有限样本界,并证明其在小样本、低维情形下表现尤为出色。同时,变分方法在高维复杂场景中更具优势,已在真实与模拟数据上的图模型估计任务中验证有效。
原文摘要 · Abstract (English)
Score matching is a vital tool for learning the distribution of data with applications across many areas including diffusion processes, energy based modelling, and graphical model estimation. Despite all these applications, little work explores its use when data is incomplete. We address this by adapting score matching (and its major extensions) to work with missing data in a flexible setting where data can be partially missing over any subset of the coordinates. We provide two separate score matching variations for general use, an importance weighting (IW) approach, and a variational approach. We provide finite sample bounds for our IW approach in finite domain settings and show it to have especially strong performance in small sample lower dimensional cases. Complementing this, we show our variational approach to be strongest in more complex high-dimensional settings which we demonstrate on graphical model estimation tasks on both real and simulated data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。