arXiv:2606.08492cs.CVcs.AI2026-06

让提示词重写与图像内容对齐,提升文生图的准确性。

Seeing is Believing: Aligning Prompt Rewriting with Visual Anchors for Text-to-Image Generation

论文配图:Seeing is Believing: Aligning Prompt Rewriting with Visual Anchors for Text-to-Image Generation
图 1 · 摘自论文原文
  • 用多模态大模型生成中间图像作为视觉锚点
  • 基于图像反馈优化提示词,生成更符合意图的内容
  • 小模型可部署,适合实际应用中的快速生成

尽管文本到图像(T2I)模型能力强大,但用户提示常因简略和模糊导致意图与生成结果之间的差距。现有方法主要提升提示的流畅性,却缺乏视觉依据,容易过度推断缺失信息,加剧偏差。为此,我们提出 FaithRewriter,一种面向 T2I 生成的新型提示增强框架。该框架首先利用多模态大语言模型(MLLM)根据原始提示生成一张中间图像作为视觉锚点,再将提示与图像联合输入大规模语言模型,生成具有视觉一致性的内容扩充;最后将这些增强结果蒸馏至小型语言模型,实现高效部署。实验表明,FaithRewriter 生成的提示在忠实度和视觉合理性上均优于强基线方法,有效缩小了意图与生成结果间的差距。

原文摘要 · Abstract (English)

Despite the impressive capabilities of text-to-image (T2I) models, an intent-generation gap often persists due to the brevity and ambiguity of user prompts. Existing approaches primarily polish the prompt for fluency and readability. However, the enhancement process still lacks visual grounding. As a result, the rewriter may over-infer missing details, causing an intent-generation gap. To address this limitation, we propose FaithRewriter, a novel prompt-enhancement framework for T2I generation. Specifically, FaithRewriter first leverages a multimodal MLLM to generate an image from the original prompt as an intermediate visual cue. This cue is then combined with the prompt and fed into a large-scale LLM to produce visually grounded augmentations that better reflect how the intended content should appear in images. Finally, these augmentations are distilled into a small-scale LLM for efficient deployment, enhancing its ability to generate effective T2I prompts. Experiments show that FaithRewriter yields prompts that are more faithful to the user intent and more visually plausible than strong baselines, helping narrow the intent-generation gap.

文生图提示工程多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。