arXiv:2602.06886cs.CV2026-02

解决文本生成图像中提示词随深度遗忘的问题

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation

  • 将早期层的提示词信息重新注入深层,缓解语义遗忘
  • 在多个数据集上提升指令遵循能力与图像质量
  • 无需训练,适用于SD3、FLUX.1等主流模型

文本到图像生成的多模态扩散变压器(MMDiTs)采用独立的文本与图像分支,并在去噪过程中实现双向信息流动。我们观察到一种提示词遗忘现象:随着网络深度增加,文本分支中的提示词表征语义逐渐丢失。通过探测三个代表性模型(SD3、SD3.5、FLUX.1)文本分支各层的语言属性,验证了该现象。为此,我们提出无需训练的提示词重注方法,将早期层的提示词表示重新注入深层以缓解遗忘。在GenEval、DPG和T2I-CompBench++上的实验表明,该方法显著提升了指令遵循能力,并在偏好、美学及整体图文生成质量指标上取得一致改进。

原文摘要 · Abstract (English)

Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information flow between text tokens and visual latents throughout denoising. In this setting, we observe a prompt forgetting phenomenon: the semantics of the prompt representation in the text branch is progressively forgotten as depth increases. We further verify this effect on three representative MMDiTs--SD3, SD3.5, and FLUX.1 by probing linguistic attributes of the representations over the layers in the text branch. Motivated by these findings, we introduce a training-free approach, prompt reinjection, which reinjects prompt representations from early layers into later layers to alleviate this forgetting. Experiments on GenEval, DPG, and T2I-CompBench++ show consistent gains in instruction-following capability, along with improvements on metrics capturing preference, aesthetics, and overall text--image generation quality.

文本生成图像扩散模型提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。