无需训练,实时修正视觉生成中的语义错误
Training-Free Semantic Correction for Autoregressive Visual Models

- 用大语言模型分析中间生成状态,识别语义偏差
- 通过回溯重生成,使结果更贴合原始提示
- 适用于多类自回归视觉模型,提升图像视频准确性
基于下一尺度预测的自回归视觉模型(AVMs)已成为图像与视频生成的重要范式。然而,将生成过程分解为不同粒度的离散尺度,使得语义错误难以发现和修正,影响最终输出质量。现有增强方法分为基于训练与无训练两类:前者计算成本高,后者忽略中间状态,导致语义误差累积。本文提出Gazer框架,在不需额外训练的前提下,将多模态大语言模型反馈融入AVM采样循环,实现生成过程中的语义修正。该框架包含两个协作阶段:反射诊断阶段从中间状态识别语义错误,语义修正阶段回溯并纠正生成轨迹,使其重新对齐目标提示。在组合图像与视频基准测试中,Gazer显著提升了多个AVMs的语义一致性和组合准确性。
原文摘要 · Abstract (English)
Autoregressive visual models (AVMs) based on next-scale prediction have emerged as a prominent paradigm for image and video synthesis. However, decomposing the generation process into discrete scales with varying granularities in AVM makes semantic errors difficult to identify and correct, thereby undermining the quality of the final output. Prior efforts to enhance AVM can be categorized into training-based and training-free approaches. Although training-based efforts to enhance AVM generation quality come at substantial computational cost, existing training-free methods neglect intermediate generation states, leaving semantic errors undiagnosed and allowing them to accumulate into the final output. In this paper, we focus on training-free paradigms and propose Gazer, a framework that integrates multimodal large language model feedback into the AVM sampling loop for in-generation semantic correction. Concretely, Gazer operates via two cooperating stages: the Reflective Diagnosis stage diagnoses semantic errors from intermediate states, while the Semantic Correction stage rewinds and rectifies the generation trajectory to realign with the target prompt. Experiments on compositional image and video benchmarks demonstrate that Gazer improves semantic alignment and compositional accuracy across multiple AVMs without additional training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。