arXiv:2504.17525cs.CV2025-04被引 1

通过选择性增强关键去噪步骤信号,提升文本图像对齐效果。

Text-to-Image Alignment in Denoising-Based Models through Step Selection

  • 在去噪后期选择性增强关键步骤的信号
  • 在扩散与流匹配模型上实现领先对齐性能
  • 适合关注生成质量与语义对齐的研究者

视觉生成AI模型常面临文本-图像对齐不足和推理能力有限的问题。本文提出一种新方法,通过在关键去噪步骤选择性增强信号,基于输入语义优化图像生成。该方法克服了早期信号修改的局限性,证明后期调整可带来更优结果。我们在扩散模型与流匹配模型上进行了大量实验,验证了该方法在生成语义对齐图像方面的有效性,达到当前最优性能。结果凸显了采样阶段合理选择对提升生成质量和整体对齐的重要作用。

原文摘要 · Abstract (English)

Visual generative AI models often encounter challenges related to text-image alignment and reasoning limitations. This paper presents a novel method for selectively enhancing the signal at critical denoising steps, optimizing image generation based on input semantics. Our approach addresses the shortcomings of early-stage signal modifications, demonstrating that adjustments made at later stages yield superior results. We conduct extensive experiments to validate the effectiveness of our method in producing semantically aligned images on Diffusion and Flow Matching model, achieving state-of-the-art performance. Our results highlight the importance of a judicious choice of sampling stage to improve performance and overall image alignment.

图像生成去噪过程语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。