通过可移动的分布学习,实现高质量透明视频生成。
Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learner
- 在隐空间和噪声空间同时调整,分离RGB与透明度分布。
- 生成视频透明效果自然,视觉质量显著优于现有方法。
- 适合需要可控透明效果的视频生成研究者使用。
生成包含透明通道的RGB-A视频具有广泛应用。然而,现有方法常因RGB与透明度信息混淆而导致质量低下。本文提出学习可移动的RGB-A分布来解决此问题:通过调整隐空间和噪声空间,将透明度分布向外偏移,同时保持RGB分布稳定,从而实现无损的透明度生成。具体而言,在VAE训练中引入感知透明度的双向扩散损失,根据似然性调整分布;在噪声空间中,调整扩散采样均值并应用高斯椭圆掩码以提供透明度引导与控制能力。此外,我们构建了一个高质量的RGB-A视频数据集。相比最先进方法,本模型在视觉质量、自然度、透明度渲染、推理便捷性和可控性方面均表现更优。开源模型已发布于:https://donghaotian123.github.io/Wan-Alpha/。
原文摘要 · Abstract (English)
Generating RGB-A videos, which include alpha channels for transparency, has wide applications. However, current methods often suffer from low quality due to confusion between RGB and alpha. In this paper, we address this problem by learning shiftable RGB-A distributions. We adjust both the latent space and noise space, shifting the alpha distribution outward while preserving the RGB distribution, thereby enabling stable transparency generation without compromising RGB quality. Specifically, for the latent space, we propose a transparency-aware bidirectional diffusion loss during VAE training, which shifts the RGB-A distribution according to likelihood. For the noise space, we propose shifting the mean of diffusion noise sampling and applying a Gaussian ellipse mask to provide transparency guidance and controllability. Additionally, we construct a high-quality RGB-A video dataset. Compared to state-of-the-art methods, our model excels in visual quality, naturalness, transparency rendering, inference convenience, and controllability. The released model is available on our website: https://donghaotian123.github.io/Wan-Alpha/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。