arXiv:2410.18615cs.CVcs.AI2024-10NeurIPS被引 8

针对公平文本生成图像中提示学习导致画质下降的问题,提出新方法提升质量。

FairQueue: Rethinking Prompt Learning for Fair Text-to-Image Generation

  • 通过分析注意力图发现提示学习会扭曲生成结构
  • 提出提示队列与注意力增强,显著改善图像质量
  • 在多个敏感属性上实现更优画质,兼顾公平性

近期,基于提示学习的方法已成为公平文本到图像生成的最先进水平。该方法利用现成参考图像为每个目标敏感属性(tSA)学习包容性提示,实现公平生成。本文首次揭示该方法会导致样本质量下降。分析表明,其训练目标——对齐学习提示与参考图像的嵌入差异——可能次优,造成提示扭曲和生成图像退化。为深入验证,我们深入分析了文本到图像模型的去噪子网络,通过跨注意力图追踪学习提示的影响,提出新颖的提示切换分析(I2H 和 H2I),并建立跨注意力图的新量化表征。结果显示,在早期去噪步骤中存在异常,导致错误全局结构持续传播,进而降低生成样本质量。基于上述洞察,我们提出两个改进思路:(i) 提示队列(Prompt Queuing)和 (ii) 注意力增强(Attention Amplification)。在多种 tSA 上的大量实验表明,所提方法在图像质量上优于当前最先进方法,同时保持竞争力的公平性表现。更多资源见 FairQueue 项目主页:https://sutd-visual-computing-group.github.io/FairQueue

原文摘要 · Abstract (English)

Recently, prompt learning has emerged as the state-of-the-art (SOTA) for fair text-to-image (T2I) generation. Specifically, this approach leverages readily available reference images to learn inclusive prompts for each target Sensitive Attribute (tSA), allowing for fair image generation. In this work, we first reveal that this prompt learning-based approach results in degraded sample quality. Our analysis shows that the approach's training objective -- which aims to align the embedding differences of learned prompts and reference images -- could be sub-optimal, resulting in distortion of the learned prompts and degraded generated images. To further substantiate this claim, as our major contribution, we deep dive into the denoising subnetwork of the T2I model to track down the effect of these learned prompts by analyzing the cross-attention maps. In our analysis, we propose a novel prompt switching analysis: I2H and H2I. Furthermore, we propose new quantitative characterization of cross-attention maps. Our analysis reveals abnormalities in the early denoising steps, perpetuating improper global structure that results in degradation in the generated samples. Building on insights from our analysis, we propose two ideas: (i) Prompt Queuing and (ii) Attention Amplification to address the quality issue. Extensive experimental results on a wide range of tSAs show that our proposed method outperforms SOTA approach's image generation quality, while achieving competitive fairness. More resources at FairQueue Project site: https://sutd-visual-computing-group.github.io/FairQueue

文本生成图像公平生成提示学习注意力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。