arXiv:2506.09740cs.CVcs.AI2025-06

提出无需训练的校准方法,提升扩散模型图文像素对齐效果

ELBO-T2IAlign: A Generic ELBO-Based Method for Calibrating Pixel-level Text-Image Alignment in Diffusion Models

  • 基于证据下界(ELBO)设计无训练校准方法
  • 在小物体、遮挡物上显著改善图文错位问题
  • 通用性强,可直接用于多种扩散模型和下游任务

扩散模型虽能生成高质量图像,但其图文对齐常因训练数据偏差出现像素级错位,尤其在小尺寸、遮挡或稀有类别物体上。本文以零样本指代图像分割为代理任务,分析现有模型的像素-文本对齐缺陷。针对此问题,提出ELBO-T2IAlign:一种基于证据下界(ELBO)的训练无关校准方法,无需额外标注、重训练或结构修改,可直接适配不同扩散模型主干。在零样本指代分割、文本引导图像编辑及组合式图像生成等任务上的大量实验表明,该方法有效提升了多类下游任务中的像素级图文对齐能力。

原文摘要 · Abstract (English)

Diffusion models excel at image generation. Recent studies have shown that these models not only generate high-quality images but also encode text-image alignment information through attention maps or loss functions. This information is valuable for various downstream tasks, including segmentation, text-guided image editing, and compositional image generation. However, current methods heavily rely on the assumption of perfect text-image alignment in diffusion models, which is not the case. In this paper, we propose using zero-shot referring image segmentation as a proxy task to evaluate the pixel-level image and class-level text alignment of popular diffusion models. We conduct an in-depth analysis of pixel-text misalignment in diffusion models from the perspective of training data bias. We find that misalignment occurs in images with small-sized, occluded, or rare object classes. Therefore, we propose ELBO-T2IAlign--a simple yet effective method to calibrate pixel-text alignment in diffusion models based on the evidence lower bound (ELBO) of likelihood. ELBO-T2IAlign is training-free and generic: it requires no additional annotations, model retraining, or architectural modifications, and it can be directly applied to different diffusion backbones. Extensive experiments on zero-shot referring image segmentation, text-guided image editing, and compositional image generation verify that the proposed calibration improves pixel-text alignment across complementary downstream tasks.

扩散模型图文对齐无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。