解释扩散模型为何能创造新图像,揭示其局部拼贴机制。
An analytic theory of creativity in convolutional diffusion models

- 发现局部性与等变性是催生创意的关键先验
- 仅用一个时变超参数即可高精度预测生成结果(中位决定系数达0.95)
- 适合研究生成模型机理或理解创造性来源的学者
我们建立了卷积扩散模型创造力的解析理论。尽管最优得分匹配理论预测模型只能复现训练数据,但实际却能生成大量原创图像。为弥合理论与实验差距,我们识别出两个简单归纳偏置:局部性和等变性,它们使模型无法达到最优得分匹配,从而产生组合式创意。由此导出完全可解析、可机械解释的局部得分(LS)和等变局部得分(ELS)机器,仅需校准一个时变超参数,即可对仅使用卷积的扩散模型(如ResNet、UNet)生成结果进行高精度定量预测(在CIFAR10、FashionMNIST、MNIST、CelebA上中位$ r^2 $分别为0.95、0.94、0.94、0.96)。理论揭示了创意源于局部补丁的多尺度、多位置拼贴机制,模型通过混合不同局部训练样本生成指数级新增图像。该理论还能部分预测预训练自注意力增强的UNet输出(在CIFAR10上中位$ r^2 \sim 0.77 $),暗示注意力在从局部拼贴中构建语义连贯性中的关键作用。
原文摘要 · Abstract (English)
We obtain an analytic, interpretable and predictive theory of creativity in convolutional diffusion models. Indeed, score-matching diffusion models can generate highly original images that lie far from their training data. However, optimal score-matching theory suggests that these models should only be able to produce memorized training examples. To reconcile this theory-experiment gap, we identify two simple inductive biases, locality and equivariance, that: (1) induce a form of combinatorial creativity by preventing optimal score-matching; (2) result in fully analytic, completely mechanistically interpretable, local score (LS) and equivariant local score (ELS) machines that, (3) after calibrating a single time-dependent hyperparameter can quantitatively predict the outputs of trained convolution only diffusion models (like ResNets and UNets) with high accuracy (median $r^2$ of $0.95, 0.94, 0.94, 0.96$ for our top model on CIFAR10, FashionMNIST, MNIST, and CelebA). Our model reveals a locally consistent patch mosaic mechanism of creativity, in which diffusion models create exponentially many novel images by mixing and matching different local training set patches at different scales and image locations. Our theory also partially predicts the outputs of pre-trained self-attention enabled UNets (median $r^2 \sim 0.77$ on CIFAR10), revealing an intriguing role for attention in carving out semantic coherence from local patch mosaics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。