用强化学习提升视觉自回归模型的主体一致性,生成更逼真的图像。
DreamVAR: Taming Reinforced Visual Autoregressive Model for High-Fidelity Subject-Driven Image Generation
- 通过预填充主体特征序列,简化自回归依赖关系。
- 在多尺度条件生成中显著提升主体保真度,优于主流扩散模型。
- 适合需要高精度主体保持的图像生成任务,如人物风格迁移。
基于扩散模型的主体驱动图像生成近期取得显著进展,展现出生成高质量图像的强大能力。然而,尽管视觉自回归(VAR)模型具备统一架构和高效推理的优势,其潜力仍未被充分挖掘。本文提出DreamVAR,一种基于VAR模型的新型主体驱动图像合成框架,采用跨尺度预测机制。技术上,先由视觉分词器提取参考主体的多尺度特征;不同于将条件特征与目标图像令牌跨尺度交错,DreamVAR在预测目标图像令牌前预先填入完整的主体特征序列。该设计简化了自回归依赖关系,并缓解了VAR范式下多尺度条件生成中的训练-测试差异。此外,DreamVAR引入强化学习,联合优化语义对齐与主体一致性。大量实验表明,相比领先扩散模型,DreamVAR在外观保留方面表现更优。
原文摘要 · Abstract (English)
Recent advances in subject-driven image generation using diffusion models have attracted considerable attention for their remarkable capabilities in producing high-quality images. Nevertheless, the potential of Visual Autoregressive (VAR) models, despite their unified architecture and efficient inference, remains underexplored. In this work, we present DreamVAR, a novel framework for subject-driven image synthesis built upon a VAR model that employs next-scale prediction. Technically, multi-scale features of the reference subject are first extracted by a visual tokenizer. Instead of interleaving these conditional features with target image tokens across scales, our DreamVAR pre-fills the full subject feature sequence prior to predicting target image tokens. This design simplifies autoregressive dependencies and mitigates the train-test discrepancy in multi-scale conditioning scenario within the VAR paradigm. DreamVAR further incorporates reinforcement learning to jointly enhance semantic alignment and subject consistency. Extensive experiments demonstrate that DreamVAR achieves superior appearance preservation compared to leading diffusion-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。