arXiv:2410.00321cs.CV2024-10NeurIPS被引 21

解决图文模型中文本嵌入的信息错乱问题,提升多物体生成准确性。

A Cat Is A Cat (Not A Dog!): Unraveling Information Mix-ups in Text-to-Image Encoders through Causal Analysis and Embedding Optimization

  • 通过因果分析揭示文本嵌入如何影响图像生成
  • 提出无需训练的优化方法,信息平衡提升125.42%
  • 新评估指标与人工判断一致率达81%,更准确衡量物体存在与精度

本文分析了文本到图像扩散模型中文本编码器中的因果机制对信息偏倚和损失的影响。以往研究主要关注去噪过程,但缺乏对文本嵌入在生成多个对象时作用的探讨。本文系统分析了文本嵌入如何影响生成结果,以及为何信息会偏向首个提及的对象。为此,提出一种无需训练的文本嵌入平衡优化方法,在Stable Diffusion上实现信息平衡提升125.42%。此外,提出一种新自动评估指标,相比现有分布评分(如CLIP文本-图像相似度)更准确量化信息损失,与人类评估一致性达81%。

原文摘要 · Abstract (English)

This paper analyzes the impact of causal manner in the text encoder of text-to-image (T2I) diffusion models, which can lead to information bias and loss. Previous works have focused on addressing the issues through the denoising process. However, there is no research discussing how text embedding contributes to T2I models, especially when generating more than one object. In this paper, we share a comprehensive analysis of text embedding: i) how text embedding contributes to the generated images and ii) why information gets lost and biases towards the first-mentioned object. Accordingly, we propose a simple but effective text embedding balance optimization method, which is training-free, with an improvement of 125.42% on information balance in stable diffusion. Furthermore, we propose a new automatic evaluation metric that quantifies information loss more accurately than existing methods, achieving 81% concordance with human assessments. This metric effectively measures the presence and accuracy of objects, addressing the limitations of current distribution scores like CLIP's text-image similarities.

图文生成嵌入优化评估指标因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。