arXiv:2608.07948cs.CV2026-08中稿 · ECCV

提升视觉自回归模型在复杂场景下的生成质量,抑制误差传播

SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model

论文配图:SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model
图 1 · 摘自论文原文
  • 引入空间-语义协同控制策略,无须训练即可增强生成能力
  • 在COCO-Stuff数据集上,生成图像的mIoU提升12.3%,细节更清晰
  • 适合需要高质量图像生成且无法重新训练模型的研究者

视觉自回归模型(VAR)因其逐尺度预测范式广受欢迎,但在处理多物体、多属性的复杂场景时面临严重性能瓶颈。现有基于扩散的方法无法有效解决VAR中跨尺度误差传播与累积问题。为此,我们提出SynVAR——首个专为VAR设计的无需训练的增强框架,引入空间-语义协同控制策略,有效抑制误差传播并提升生成质量。SynVAR包含三个核心组件:(1) 全局引导,在早期确保合理的空间结构;(2) 感受野约束,缓解早期语义混淆;(3) 高频补偿,恢复细粒度细节。大量定量与定性实验表明,SynVAR显著提升了VAR在复杂场景建模上的能力。

原文摘要 · Abstract (English)

VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling complex scenes with multiple objects and attributes. Existing diffusion-based enhancement methods fail to adequately address the unique challenge of cross-scale error propagation and accumulation in VAR. To this end, we propose SynVAR, the first training-free enhancement framework specifically tailored for the VAR paradigm, which introduces a spatial-semantic collaborative control strategy to effectively suppress propagation error and improve generation quality. SynVAR comprises three key components: (1) Global guidance to ensure reasonable spatial structure in the early stages, (2) Receptive field constraints to mitigate early-stage semantic confusion, (3) High-frequency compensation to recover fine-grained details. Extensive quantitative and qualitative experiments demonstrate the significant improvements in the ability of SynVAR to enhance the VAR's capability for complex scene modeling.

视觉生成自回归模型误差抑制图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。