用状态空间模型提升图像风格迁移的效率与质量
MambaStyle: Efficient StyleGAN Inversion for Real Image Editing with State-Space Models
- 采用视觉状态空间模型构建单阶段编码器,实现高效图像反演
- 在保持高重建质量的同时,参数量和计算量显著低于现有方法
- 适合需要实时编辑的场景,如交互式图像处理
将真实图像反演到StyleGAN潜在空间以实现属性编辑的任务已被广泛研究。然而,现有GAN反演方法难以在高重建质量、有效编辑性和计算效率之间取得平衡。本文提出MambaStyle,一种基于视觉状态空间模型(VSSMs)的单阶段编码器方法,用于GAN反演与编辑。该方法通过在架构中引入VSSMs,实现了高质量图像反演与灵活编辑,同时显著减少参数量和计算复杂度。大量实验表明,MambaStyle在反演精度、编辑质量与计算效率之间达到了更优平衡。值得注意的是,本方法在降低模型复杂度和加快推理速度的同时,仍能获得更优的反演与编辑效果,适用于实时应用。
原文摘要 · Abstract (English)
The task of inverting real images into StyleGAN's latent space to manipulate their attributes has been extensively studied. However, existing GAN inversion methods struggle to balance high reconstruction quality, effective editability, and computational efficiency. In this paper, we introduce MambaStyle, an efficient single-stage encoder-based approach for GAN inversion and editing that leverages vision state-space models (VSSMs) to address these challenges. Specifically, our approach integrates VSSMs within the proposed architecture, enabling high-quality image inversion and flexible editing with significantly fewer parameters and reduced computational complexity compared to state-of-the-art methods. Extensive experiments show that MambaStyle achieves a superior balance among inversion accuracy, editing quality, and computational efficiency. Notably, our method achieves superior inversion and editing results with reduced model complexity and faster inference, making it suitable for real-time applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。