用内在奖励提升扩散模型对复杂提示的生成一致性。
TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward

- 基于概念分布重叠定义内在奖励,无需外部监督。
- 在T2ICompBench上提升组合一致性,图像质量不变。
- 适合需要精准提示理解的文本生成图像场景。
近期强大的文本到图像生成模型促使测试时方法的发展,以调整采样轨迹,使图像更忠实于复杂的组合提示。我们提出TILT,一种无需训练的测试时奖励对齐框架,用于组合式文本到图像生成。我们将组合失败解释为联合概念与单概念分布之间的重叠模式,并定义一个奖励函数,鼓励所有概念同时出现的样本。该奖励内生于基础模型,无需任何外部监督或奖励模型。由此得到一个带有闭式倾斜目标分布的KL约束优化目标,以及指导扩散采样的合理步骤。概念分布间的相互作用结合上述奖励,自然引出两种不同的引导策略,而一种平衡二者优势的混合方法表现更强。在T2ICompBench提示上的实验表明,相比先前基线,本方法在保持图像质量的同时提升了组合对齐度。
原文摘要 · Abstract (English)
Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present TILT, a training-free framework for compositional text-to-image generation via test-time reward alignment. We interpret compositional failures as overlap modes between joint and single-concept distributions, and define a reward that favors samples where all concepts are jointly present. This reward is intrinsic to the base model and does not require any external supervision or reward models. This yields a KL-constrained objective with a closed-form tilted target distribution and principled guiding steps for diffusion sampling. The interaction of concept distributions together with the above reward naturally leads to two different guidance strategies while a hybrid approach that balances their respective benefits produces stronger performance. Experiments on prompts from T2ICompBench show that our method improves compositional alignment while preserving image quality compared to previous baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。