arXiv:2603.28367cs.CV2026-03被引 1

提升文本引导图像编辑的结构一致性,速度更快更精准。

Rethinking Structure Preservation in Text-Guided Image Editing with Visual Autoregressive Models

  • 分步定位可编辑区域,平衡修改精度与背景保留。
  • 通过特征注入增强编辑后图像的结构连贯性。
  • 自适应调整注入比例,适合复杂图像修改场景。

视觉自回归(VAR)模型作为新兴生成模型,在文本引导图像编辑中展现出潜力,将编辑范式从扩散模型的噪声操作转向像素级令牌操作,实现更好的背景保留和显著更快的推理速度。然而,现有方法仍面临两大挑战:精确识别可编辑令牌位置、保持编辑结果的结构一致性。本文通过分析VAR模型中间特征分布,提出新框架:首先设计粗到精的令牌定位策略,优化可编辑区域;其次分析中间表示,识别结构相关特征,并构建简单有效的特征注入机制以增强编辑前后图像的结构一致性;最后提出基于强化学习的自适应特征注入方案,自动学习不同尺度与层的注入比例,协同优化编辑保真度与结构保留。大量实验表明,该方法在局部与全局编辑场景下均优于现有最先进方法,显著提升结构一致性和编辑质量。

原文摘要 · Abstract (English)

Visual autoregressive (VAR) models have recently emerged as a promising family of generative models, enabling a wide range of downstream vision tasks such as text-guided image editing. By shifting the editing paradigm from noise manipulation in diffusion-based methods to token-level operations, VAR-based approaches achieve better background preservation and significantly faster inference. However, existing VAR-based editing methods still face two key challenges: accurately localizing editable tokens and maintaining structural consistency in the edited results. In this work, we propose a novel text-guided image editing framework rooted in an analysis of intermediate feature distributions within VAR models. First, we introduce a coarse-to-fine token localization strategy that can refine editable regions, balancing editing fidelity and background preservation. Second, we analyze the intermediate representations of VAR models and identify structure-related features, by which we design a simple yet effective feature injection mechanism to enhance structural consistency between the edited and source images. Third, we develop a reinforcement learning-based adaptive feature injection scheme that automatically learns scale- and layer-specific injection ratios to jointly optimize editing fidelity and structure preservation. Extensive experiments demonstrate that our method achieves superior structural consistency and editing quality compared with state-of-the-art approaches, across both local and global editing scenarios.

图像编辑自回归模型结构保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。