arXiv:2509.01984cs.CV2025-09被引 6

提出首个针对视觉自回归模型的文本编辑噪声逆解方法,无需训练即可精准修改图像。

Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing

  • 设计位置感知的反向采样函数,生成可逆的Gumbel噪声以还原图像。
  • 在不破坏原图背景和结构的前提下,实现符合文本提示的可控编辑。
  • 适用于需快速、无训练图像修改的场景,如内容创作与交互设计。

视觉自回归模型(VAR)近年来成为一类有前景的生成模型,在文本到图像生成任务中表现已接近扩散模型。尽管条件生成已被广泛研究,但无需额外训练即可实现提示引导的图像编辑同样至关重要,可支持众多实际应用。本文提出面向VAR模型的首个基于噪声逆解的编辑方法——视觉自回归逆噪声(VARIN)。VARIN引入一种新型伪逆函数,即位置感知最大值逆推(LAI),用于argmax采样,生成逆Gumbel噪声。这些逆噪声能精确重建源图像,并实现与文本提示对齐的精准、可控编辑。大量实验表明,VARIN可在不显著改变原始背景和结构细节的前提下,有效根据指定提示修改源图像,验证了其作为实用编辑方法的有效性。

原文摘要 · Abstract (English)

Visual autoregressive models (VAR) have recently emerged as a promising class of generative models, achieving performance comparable to diffusion models in text-to-image generation tasks. While conditional generation has been widely explored, the ability to perform prompt-guided image editing without additional training is equally critical, as it supports numerous practical real-world applications. This paper investigates the text-to-image editing capabilities of VAR by introducing Visual AutoRegressive Inverse Noise (VARIN), the first noise inversion-based editing technique designed explicitly for VAR models. VARIN leverages a novel pseudo-inverse function for argmax sampling, named Location-aware Argmax Inversion (LAI), to generate inverse Gumbel noises. These inverse noises enable precise reconstruction of the source image and facilitate targeted, controllable edits aligned with textual prompts. Extensive experiments demonstrate that VARIN effectively modifies source images according to specified prompts while significantly preserving the original background and structural details, thus validating its efficacy as a practical editing approach.

图像编辑自回归模型噪声逆解文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。