arXiv:2603.19158cs.CV2026-03中稿 · CVPR

解决罕见概念生成中图像失真问题,让扩散模型更忠实地遵循用户提示。

Adaptive Auxiliary Prompt Blending for Target-Faithful Diffusion Generation

  • 通过自适应融合辅助提示,动态平衡主提示与辅助提示的影响。
  • 在RareBench和FlowEdit数据集上,语义准确率和结构保真度显著提升。
  • 无需训练,基于数学原理实现稳定生成,适合复杂文本到图像任务。

基于扩散的文本到图像(T2I)模型在生成逼真且语义丰富的图像方面取得了显著进展。然而,当目标概念位于训练分布的低密度区域时,这些模型常产生语义错位或结构不一致的结果。这一局限源于文本-图像数据集的长尾特性,即稀有概念或编辑指令代表性不足。为此,我们提出自适应辅助提示融合(AAPB)——一种统一框架,用于稳定低密度区域的扩散过程。AAPB利用辅助锚点提示,在稀有概念生成中提供语义支持,在图像编辑中提供结构支持,确保对目标提示的忠实引导。不同于以往的启发式提示交替方法,AAPB推导出闭式自适应系数,可在每个扩散步骤中最优平衡辅助锚点与目标提示的影响。基于Tweedie恒等式,我们的公式提供了原则性且无需训练的自适应提示融合框架,确保生成稳定且目标忠实。通过受控实验验证了自适应插值优于固定插值,并在RareBench和FlowEdit数据集上实证展示了持续改进,相比先前的无训练基线,在语义准确性和结构保真度上表现更优。

原文摘要 · Abstract (English)

Diffusion-based text-to-image (T2I) models have made remarkable progress in generating photorealistic and semantically rich images. However, when the target concepts lie in low-density regions of the training distribution, these models often produce semantically misaligned or structurally inconsistent results. This limitation arises from the long-tailed nature of text-image datasets, where rare concepts or editing instructions are underrepresented. To address this, we introduce Adaptive Auxiliary Prompt Blending (AAPB) - a unified framework that stabilizes the diffusion process in low-density regions. AAPB leverages auxiliary anchor prompts to provide semantic support in rare concept generation and structural support in image editing, ensuring faithful guidance toward the target prompt. Unlike prior heuristic prompt alternation methods, AAPB derives a closed-form adaptive coefficient that optimally balances the influence between the auxiliary anchor and the target prompt at each diffusion step. Grounded in Tweedie's identity, our formulation provides a principled and training-free framework for adaptive prompt blending, ensuring stable and target-faithful generation. We demonstrate the effectiveness of adaptive interpolation over fixed interpolation through controlled experiments and empirically show consistent improvements on the RareBench and FlowEdit datasets, achieving superior semantic accuracy and structural fidelity compared to prior training-free baselines.

扩散模型图像生成提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。