提出一种新正则化方法,让文生图模型的隐空间更接近标准高斯分布。
Moment- and Power-Spectrum-Based Gaussianity Regularization for Text-to-Image Models
- 结合空间域矩匹配与频域功率谱正则,统一现有高斯性约束
- 在测试时优化中提升图像美学与文本对齐,加速收敛并防奖励劫持
- 适用于需要隐空间优化的下游任务,如生成调优与风格控制
我们提出一种新型正则化损失,强制样本服从标准高斯分布,从而促进文本到图像模型隐空间中的各类下游优化任务。将高维样本元素视为一维标准高斯变量,构建融合空间域矩正则与频域功率谱正则的复合损失。由于矩和功率谱分布的期望值可解析计算,该损失能有效引导样本符合这些统计特性。为保证置换不变性,损失作用于随机置换输入。值得注意的是,现有基于高斯性的正则化方法均属于本统一框架:部分对应特定阶数的矩损失,而此前的协方差匹配损失等价于我们的频域损失,但因在空间域计算导致更高时间复杂度。我们在测试时奖励对齐任务中展示了该正则化的应用,显著提升了图像美学与文本对齐效果,优于以往高斯性正则,有效防止奖励劫持并加速收敛。
原文摘要 · Abstract (English)
We propose a novel regularization loss that enforces standard Gaussianity, encouraging samples to align with a standard Gaussian distribution. This facilitates a range of downstream tasks involving optimization in the latent space of text-to-image models. We treat elements of a high-dimensional sample as one-dimensional standard Gaussian variables and define a composite loss that combines moment-based regularization in the spatial domain with power spectrum-based regularization in the spectral domain. Since the expected values of moments and power spectrum distributions are analytically known, the loss promotes conformity to these properties. To ensure permutation invariance, the losses are applied to randomly permuted inputs. Notably, existing Gaussianity-based regularizations fall within our unified framework: some correspond to moment losses of specific orders, while the previous covariance-matching loss is equivalent to our spectral loss but incurs higher time complexity due to its spatial-domain computation. We showcase the application of our regularization in generative modeling for test-time reward alignment with a text-to-image model, specifically to enhance aesthetics and text alignment. Our regularization outperforms previous Gaussianity regularization, effectively prevents reward hacking and accelerates convergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。