arXiv:2510.00993cs.CV2025-10EMNLP被引 1

提升自回归模型的视觉生成质量,解决序列生成中的误差累积问题。

Visual Self-Refinement for Autoregressive Models

  • 引入可插拔后训练模块,统一优化生成序列中所有视觉标记
  • 利用全局上下文关系,显著改善语义一致性与空间对应关系
  • 适用于需要高质量图像生成的视觉语言任务,如图文生成

自回归模型在视觉-语言数据建模中表现优异,但视觉信号的空间特性与逐词预测的序列依赖存在冲突,导致生成效果受限。本文提出一种即插即用的后训练精修模块,用于增强自回归模型生成序列中的复杂空间对应关系。该模块作为预训练后的补充步骤,联合优化所有生成标记,在共享的序列预测框架下提升视觉-语言建模能力。通过利用标记间的全局上下文与相互关系,有效缓解了序列生成过程中的误差累积问题。实验表明,该方法显著提升了生成质量,增强了模型生成语义一致结果的能力。

原文摘要 · Abstract (English)

Autoregressive models excel in sequential modeling and have proven to be effective for vision-language data. However, the spatial nature of visual signals conflicts with the sequential dependencies of next-token prediction, leading to suboptimal results. This work proposes a plug-and-play refinement module to enhance the complex spatial correspondence modeling within the generated visual sequence. This module operates as a post-pretraining step to jointly refine all generated tokens of autoregressive model, enhancing vision-language modeling under a shared sequential prediction framework. By leveraging global context and relationship across the tokens, our method mitigates the error accumulation issue within the sequential generation. Experiments demonstrate that the proposed method improves the generation quality, enhancing the model's ability to produce semantically consistent results.

自回归模型视觉生成序列优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。