不靠训练的简单模型也能生成逼真图像,揭示了自然图像的内在结构。
Scaling Non-Parametric Sampling with Representation
- 基于局部上下文定义像素分布,无参数、无需训练
- 在MNIST上生成高保真图像,在CIFAR-10上生成视觉逼真的结果
- 可解释性强,揭示了图像生成中的组合式泛化机制
规模化和架构进步带来了令人惊叹的逼真图像生成模型,但其内部机制仍不透明。本文目标是摒弃复杂的工程技巧,提出一种简单、非参数化的生成模型。设计基于自然图像的三个原则:(i)空间非平稳性,(ii)低层规律性,(iii)高层语义,并从局部上下文窗口定义每个像素的分布。尽管架构极简且无需训练,该模型在MNIST上生成高保真样本,在CIFAR-10上生成视觉上引人注目的图像。这种简洁性与强大性能的结合,指向自然图像结构的最小理论。模型的白盒特性使其可进行机制分析,通过追踪每个生成像素的来源,发现了一种简单的、组合式的“部分-整体泛化”过程,为大型神经网络生成模型如何学习泛化提出了假设。
原文摘要 · Abstract (English)
Scaling and architectural advances have produced strikingly photorealistic image generative models, yet their mechanisms still remain opaque. Rather than advancing scaling, our goal is to strip away complicated engineering tricks and propose a simple, non-parametric generative model. Our design is grounded in three principles of natural images-(i) spatial non-stationarity, (ii) low-level regularities, and (iii) high-level semantics-and defines each pixel's distribution from its local context window. Despite its minimal architecture and no training, the model produces high-fidelity samples on MNIST and visually compelling CIFAR-10 images. This combination of simplicity and strong empirical performance points toward a minimal theory of natural-image structure. The model's white-box nature also allows us to have a mechanistic understanding of how the model generalizes and generates diverse images. We study it by tracing each generated pixel back to its source images. These analyses reveal a simple, compositional procedure for "part-whole generalization", suggesting a hypothesis for how large neural network generative models learn to generalize.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。