arXiv:2510.12975cs.LGstat.ML2025-10中稿 · NeurIPS被引 4

用去噪得分匹配损失估算数据局部内在维数,更高效且省资源。

A Connection Between Score Matching and Local Intrinsic Dimension

  • 利用去噪得分匹配损失作为局部内在维数的下界来估算
  • 在曼达巴和Stable Diffusion 3.5上表现优于传统方法
  • 无需大量前向传播或梯度计算,适合低资源场景

数据的局部内在维数(LID)是信号处理与学习理论中的基础概念,但对高维复杂数据的LID量化长期面临挑战。近期研究发现,扩散模型通过其得分估计的谱结构及密度估计在噪声扰动下的变化率可捕捉数据的LID。然而这些方法需多次前向传播或梯度计算,限制了其在计算与内存受限场景的应用。本文证明LID是去噪得分匹配损失的下界,从而提出以该损失作为LID估计算法。进一步表明等价的隐式得分匹配损失可通过正则维度近似LID,且与最近的LID估计算法FLIPD密切相关。在曼达巴基准和Stable Diffusion 3.5上的实验表明,去噪得分匹配损失是一种极具竞争力且可扩展的LID估计算法,在问题规模增大和量化水平提高时仍保持优异准确率与极低内存开销。

原文摘要 · Abstract (English)

The local intrinsic dimension (LID) of data is a fundamental quantity in signal processing and learning theory, but quantifying the LID of high-dimensional, complex data has been a historically challenging task. Recent works have discovered that diffusion models capture the LID of data through the spectra of their score estimates and through the rate of change of their density estimates under various noise perturbations. While these methods can accurately quantify LID, they require either many forward passes of the diffusion model or use of gradient computation, limiting their applicability in compute- and memory-constrained scenarios. We show that the LID is a lower bound on the denoising score matching loss, motivating use of the denoising score matching loss as a LID estimator. Moreover, we show that the equivalent implicit score matching loss also approximates LID via the normal dimension and is closely related to a recent LID estimator, FLIPD. Our experiments on a manifold benchmark and with Stable Diffusion 3.5 indicate that the denoising score matching loss is a highly competitive and scalable LID estimator, achieving superior accuracy and memory footprint under increasing problem size and quantization level.

局部维数扩散模型得分匹配高效估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。