arXiv:2506.01929cs.GRcs.AI2025-06被引 10

让图像生成模型读懂前后矛盾的提示词,准确还原语义。

Image Generation from Contextually-Contradictory Prompts

  • 分阶段分解提示词,用代理提示引导生成过程
  • 在多个测试提示上显著提升与文本的匹配度
  • 适合处理带有隐含冲突的复杂描述场景

文本到图像的扩散模型在自然语言提示下能生成高质量、多样化的图像。然而,当提示包含与其学习先验相矛盾的概念组合时,往往无法生成语义准确的结果。我们定义这种失败模式为上下文矛盾,即一个概念因训练中习得的纠缠关联而隐式否定另一个。为此,我们提出一种阶段感知的提示分解框架,通过一系列代理提示引导去噪过程。每个代理提示根据去噪特定阶段预期出现的语义内容构建,同时保证上下文一致性。通过大语言模型分析目标提示,识别矛盾并生成保留原意但消除冲突的替代表达,实现提示信息与去噪进程的对齐。实验表明,在多种具有挑战性的提示下,该方法显著提升了生成结果与文本提示的一致性。

原文摘要 · Abstract (English)

Text-to-image diffusion models excel at generating high-quality, diverse images from natural language prompts. However, they often fail to produce semantically accurate results when the prompt contains concept combinations that contradict their learned priors. We define this failure mode as contextual contradiction, where one concept implicitly negates another due to entangled associations learned during training. To address this, we propose a stage-aware prompt decomposition framework that guides the denoising process using a sequence of proxy prompts. Each proxy prompt is constructed to match the semantic content expected to emerge at a specific stage of denoising, while ensuring contextual coherence. To construct these proxy prompts, we leverage a large language model (LLM) to analyze the target prompt, identify contradictions, and generate alternative expressions that preserve the original intent while resolving contextual conflicts. By aligning prompt information with the denoising progression, our method enables fine-grained semantic control and accurate image generation in the presence of contextual contradictions. Experiments across a variety of challenging prompts show substantial improvements in alignment to the textual prompt.

图像生成文本控制提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。