arXiv:2502.16965cs.CV2025-02被引 1

用视觉全貌提示提升自回归图像生成的结构准确性和稳定性

Autoregressive Image Generation with Vision Full-view Prompt

  • 设计图像专用全貌提示,模拟人类先看整体再生成的视觉逻辑
  • 相比无提示方法,生成质量提升约20%,推理步骤更稳定
  • 适合需要高质量结构化图像生成的研究者与应用开发

在自回归(AR)图像生成中,基于大语言模型‘预测下一个标记’范式的模型已达到与扩散模型相当的性能,但直接套用语言模型处理复杂图像时,难以还原图像结构与细节,影响生成准确性和稳定性。此外,该范式与人类视觉感知中的上下文扫描和逻辑推理过程不匹配,限制了生成效果。受自然语言处理中提示工程的启发,本文提出视觉全貌提示(VF prompt),专为自回归图像生成设计,模拟人类先感知整体分布再逐步生成的创作过程。通过引入全貌提示,模型获得更强的上下文逻辑能力,并在增加推理步数后显著提升生成稳定性。实验表明,相比无提示的AR方法,本方法性能提升约20%。

原文摘要 · Abstract (English)

In autoregressive (AR) image generation, models based on the 'next-token prediction' paradigm of LLMs have shown comparable performance to diffusion models by reducing inductive biases. However, directly applying LLMs to complex image generation can struggle with reconstructing the image's structure and details, impacting the generation's accuracy and stability. Additionally, the 'next-token prediction' paradigm in the AR model does not align with the contextual scanning and logical reasoning processes involved in human visual perception, limiting effective image generation. Prompt engineering, as a key technique for guiding LLMs, leverages specifically designed prompts to improve model performance on complex natural language processing (NLP) tasks, enhancing accuracy and stability of generation while maintaining contextual coherence and logical consistency, similar to human reasoning. Inspired by prompt engineering from the field of NLP, we propose Vision Full-view prompt (VF prompt) to enhance autoregressive image generation. Specifically, we design specialized image-related VF prompts for AR image generation to simulate the process of human image creation. This enhances contextual logic ability by allowing the model to first perceive overall distribution information before generating the image, and improve generation stability by increasing the inference steps. Compared to the AR method without VF prompts, our method shows outstanding performance and achieves an approximate improvement of 20%.

自回归生成视觉提示图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。