arXiv:2601.21419cs.LGcs.CV2026-01被引 5

提出新理论解释为何高维数据生成应直接预测数据本身

Revisiting Diffusion Model Predictions Through Dimensionality

  • 构建通用预测框架,统一噪声、速度与数据预测
  • 理论证明当数据内在维度远小于环境维度时,直接预测数据最优
  • 设计k-Diff自动学习最佳预测目标,无需估算数据维度

近期扩散与流匹配模型的进展表明,高维场景下最优预测目标正从噪声(ε)和速度(v)转向直接预测数据(x)。然而,为何最优目标依赖于数据具体特性尚无明确解释。本文基于广义预测形式化框架,涵盖任意输出目标,其中ε、v、x预测均为特例。推导出数据几何结构与最优预测目标之间的解析关系,严格证明当环境维度显著高于数据内在维度时,x预测更优。尽管理论指出维度是决定因素,但实际中数据流形的内在维度通常难以估计。为此,我们提出k-Diff框架,采用数据驱动方法直接从数据中学习最优预测参数k,避免显式维度估计。在潜在空间与像素空间图像生成中的大量实验表明,k-Diff在不同架构与数据规模下均持续优于固定目标基线,提供一种原理严谨且自动化的生成性能提升方法。

原文摘要 · Abstract (English)

Recent advances in diffusion and flow matching models have highlighted a shift in the preferred prediction target -- moving from noise ($\varepsilon$) and velocity (v) to direct data (x) prediction -- particularly in high-dimensional settings. However, a formal explanation of why the optimal target depends on the specific properties of the data remains elusive. In this work, we provide a theoretical framework based on a generalized prediction formulation that accommodates arbitrary output targets, of which $\varepsilon$-, v-, and x-prediction are special cases. We derive the analytical relationship between data's geometry and the optimal prediction target, offering a rigorous justification for why x-prediction becomes superior when the ambient dimension significantly exceeds the data's intrinsic dimension. Furthermore, while our theory identifies dimensionality as the governing factor for the optimal prediction target, the intrinsic dimension of manifold-bound data is typically intractable to estimate in practice. To bridge this gap, we propose k-Diff, a framework that employs a data-driven approach to learn the optimal prediction parameter k directly from data, bypassing the need for explicit dimension estimation. Extensive experiments in both latent-space and pixel-space image generation demonstrate that k-Diff consistently outperforms fixed-target baselines across varying architectures and data scales, providing a principled and automated approach to enhancing generative performance.

扩散模型生成模型维度分析自适应预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。