通过中间文本表示增强图像生成,让唯一对象更忠实于文字描述。
Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment

- 在扩散过程早期注入文本编码器的中间隐藏状态,恢复被忽略的概念信息。
- 在唯一对象上提升19.1个百分点的VQAScore,同时保持生成质量与人类偏好。
- 专为挑战性提示设计评测基准,适合需要精准图文对齐的研究者。
文本到图像(T2I)扩散模型常因概念关联偏差而无法忠实还原显式文字描述,尤其在唯一对象(OAO)如天体、地标、艺术品等上表现更差,其固有视觉身份难以通过提示调整。我们通过信息论分析发现,最终文本嵌入会丢失中间层中的概念级信息,削弱后续去噪过程的互信息。为此提出中间文本表示(IR)引导的扩散方法,无需额外训练或模型,在早期去噪步骤中注入文本编码器的中间状态,恢复被抑制的概念。为系统评估该任务,我们构建了OAO-AttackBench基准,包含直接违背唯一对象核心视觉身份的反事实提示。在四个基准(包括OAO-AttackBench)上的实验表明,该方法在保持生成保真度和人类偏好前提下,使VQAScore最高提升19.1个百分点。
原文摘要 · Abstract (English)
Text-to-image (T2I) diffusion models often fail to faithfully render explicit textual descriptions, instead defaulting to strongly learned visual priors due to a phenomenon referred to as concept association bias. We show that such bias is particularly strong for one-and-only (OAO) objects, entities that exist in a single canonical form, such as celestial bodies, landmarks, and artworks. The deeply ingrained visual identity for these concepts often resists modification through prompting alone. Addressing this challenge, we first identify through an information-theoretic analysis that the final text embedding discards concept-level information present in the intermediate-layer text representations, reducing the mutual information available to the subsequent denoising process. We then propose Intermediate Text Representation (IR)-guided diffusion, which injects intermediate hidden states of the text encoder into the conditioning signal during early denoising steps, recovering suppressed concepts without any additional training, optimization, or external models. To systematically evaluate the challenging task of aligning generative outputs with unusual prompts for OAO objects, we introduce OAO-AttackBench, a benchmark comprising counterfactual prompts that directly conflict with the core visual identity of OAO objects. Experiments on four benchmarks, including OAO-AttackBench, show that our method achieves up to a 19.1 percentage-point improvement in VQAScore while preserving generation fidelity and human preference. Project page: https://soyoun-won.github.io/one-and-only-ir-guidance/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。