arXiv:2506.00849cs.LGcs.AI2025-06ICLR被引 2

统一分析生成模型泛化能力,揭示扩散时间对性能的关键影响

Generalization in VAE and Diffusion Models: A Unified Information-Theoretic Analysis

  • 将编码器与生成器视为随机映射,构建统一理论框架
  • 发现扩散模型泛化存在随扩散时间T变化的明确权衡关系
  • 给出仅依赖训练数据的可计算边界,可用于优化模型超参数

尽管扩散模型(DMs)和变分自编码器(VAEs)在实践中表现良好,但其泛化性能仍缺乏理论探讨,尤其未充分考虑共享编码-生成结构的影响。本文利用最新的信息论工具,提出一个统一的理论框架,将编码器与生成器视为随机映射,为两者泛化提供保证。该框架实现:(1) 对VAE进行精细化分析,首次纳入生成器的泛化性;(2) 揭示扩散模型泛化性能与扩散时间$T$之间的显式权衡;(3) 提出仅基于训练数据的可计算泛化边界,支持最优$T$选择,并可嵌入优化过程以提升性能。在合成与真实数据集上的实验验证了理论的有效性。

原文摘要 · Abstract (English)

Despite the empirical success of Diffusion Models (DMs) and Variational Autoencoders (VAEs), their generalization performance remains theoretically underexplored, especially lacking a full consideration of the shared encoder-generator structure. Leveraging recent information-theoretic tools, we propose a unified theoretical framework that provides guarantees for the generalization of both the encoder and generator by treating them as randomized mappings. This framework further enables (1) a refined analysis for VAEs, accounting for the generator's generalization, which was previously overlooked; (2) illustrating an explicit trade-off in generalization terms for DMs that depends on the diffusion time $T$; and (3) providing computable bounds for DMs based solely on the training data, allowing the selection of the optimal $T$ and the integration of such bounds into the optimization process to improve model performance. Empirical results on both synthetic and real datasets illustrate the validity of the proposed theory.

生成模型泛化分析信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。