提出CSD-VAR模型,实现视觉自回归模型中的内容与风格解耦。
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
- 基于尺度感知的交替优化,提升内容与风格表征分离度。
- 使用SVD修正方法,减少内容信息在风格表示中的泄露。
- 适合需要高保真内容保留与灵活风格迁移的应用场景。
将单张图像中的内容与风格解耦(即内容-风格分解,CSD),可实现内容再语境化与风格迁移,提升视觉合成的创作灵活性。尽管近期个性化方法已探索显式内容风格的分解,但多针对扩散模型。视觉自回归建模(VAR)作为新兴生成框架,采用逐尺度预测范式,性能已达扩散模型水平。本文首次将VAR用于CSD任务,利用其尺度级生成过程提升解耦效果。提出CSD-VAR,引入三项创新:(1) 尺度感知交替优化策略,使内容与风格表征与其对应尺度对齐;(2) 基于SVD的修正方法,缓解内容泄露至风格表征的问题;(3) 增强型键值(K-V)记忆模块,强化内容身份保持。为评估该任务,构建了专用于CSD的CSD-100数据集,包含多样主体以多种艺术风格呈现。实验表明,CSD-VAR优于现有方法,在内容保留和风格还原精度上均取得更优表现。
原文摘要 · Abstract (English)
Disentangling content and style from a single image, known as content-style decomposition (CSD), enables recontextualization of extracted content and stylization of extracted styles, offering greater creative flexibility in visual synthesis. While recent personalization methods have explored the decomposition of explicit content style, they remain tailored for diffusion models. Meanwhile, Visual Autoregressive Modeling (VAR) has emerged as a promising alternative with a next-scale prediction paradigm, achieving performance comparable to that of diffusion models. In this paper, we explore VAR as a generative framework for CSD, leveraging its scale-wise generation process for improved disentanglement. To this end, we propose CSD-VAR, a novel method that introduces three key innovations: (1) a scale-aware alternating optimization strategy that aligns content and style representation with their respective scales to enhance separation, (2) an SVD-based rectification method to mitigate content leakage into style representations, and (3) an Augmented Key-Value (K-V) memory enhancing content identity preservation. To benchmark this task, we introduce CSD-100, a dataset specifically designed for content-style decomposition, featuring diverse subjects rendered in various artistic styles. Experiments demonstrate that CSD-VAR outperforms prior approaches, achieving superior content preservation and stylization fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。