arXiv:2505.03203cs.CV2025-05

通过优化噪声选择与精准掩码控制,提升复杂文本提示下的图文对齐效果。

PiCo: Enhancing Text-Image Alignment with Improved Noise Selection and Precise Mask Control in Diffusion Models

  • 基于质量评估筛选初始噪声,提升生成起点适配性。
  • 引入像素级引用掩码,精确调控跨注意力交互。
  • 无需训练,可直接用于现有扩散模型,适合复杂图文生成场景。

先进扩散模型在文本到图像的组合生成上取得了显著进展,但在面对复杂文本提示时,仍难以实现良好的图文对齐。本文指出两个关键影响因素:随机初始化噪声的质量和生成控制掩码的可靠性。为此提出PiCo(Pick-and-Control)——一种无需训练的新方法,包含两个核心组件:首先,设计噪声选择模块,评估噪声质量并判断其是否适配目标文本,采用快速采样策略保证效率;其次,引入引用掩码模块,生成像素级掩码,精确调节跨注意力图。该掩码被融入标准扩散过程,引导文本与图像特征的合理交互。大量实验验证了PiCo在减少用户手动调参负担、提升多样文本描述下图文对齐能力方面的有效性。

原文摘要 · Abstract (English)

Advanced diffusion models have made notable progress in text-to-image compositional generation. However, it is still a challenge for existing models to achieve text-image alignment when confronted with complex text prompts. In this work, we highlight two factors that affect this alignment: the quality of the randomly initialized noise and the reliability of the generated controlling mask. We then propose PiCo (Pick-and-Control), a novel training-free approach with two key components to tackle these two factors. First, we develop a noise selection module to assess the quality of the random noise and determine whether the noise is suitable for the target text. A fast sampling strategy is utilized to ensure efficiency in the noise selection stage. Second, we introduce a referring mask module to generate pixel-level masks and to precisely modulate the cross-attention maps. The referring mask is applied to the standard diffusion process to guide the reasonable interaction between text and image features. Extensive experiments have been conducted to verify the effectiveness of PiCo in liberating users from the tedious process of random generation and in enhancing the text-image alignment for diverse text descriptions.

扩散模型图文对齐生成控制零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。