arXiv:2501.03006cs.CV2025-01CVPR被引 15

让文本生成视频支持透明度,实现烟雾、反光等特效无缝融合。

TransPixeler: Advancing Text-to-Video Generation with Transparency

论文配图:TransPixeler: Advancing Text-to-Video Generation with Transparency
图 1 · 摘自论文原文
  • 用特定令牌和LoRA微调,让模型同时生成彩色与透明通道。
  • 在数据有限下仍保持颜色与透明度高度一致。
  • 适合影视特效、交互内容创作人员使用。

文本到视频生成模型已取得显著进展,广泛应用于娱乐、广告和教育。然而,生成包含透明度通道(RGBA)的视频仍面临挑战,主要受限于数据集稀缺及现有模型难以适配。透明通道对视觉特效(VFX)至关重要,可实现烟雾、反光等元素与场景自然融合。我们提出TransPixeler,一种扩展预训练视频模型以生成RGBA视频的方法,同时保留原始RGB能力。该方法基于扩散变换器(DiT)架构,引入针对透明度的专用令牌,并采用LoRA微调,联合生成高一致性的RGB与alpha通道。通过优化注意力机制,模型在有限训练数据下仍能保持原有RGB模型的优势,实现颜色与透明度的良好对齐。所提方法有效生成多样且一致的RGBA视频,推动了视觉特效与交互内容创作的发展。

原文摘要 · Abstract (English)

Text-to-video generative models have made significant strides, enabling diverse applications in entertainment, advertising, and education. However, generating RGBA video, which includes alpha channels for transparency, remains a challenge due to limited datasets and the difficulty of adapting existing models. Alpha channels are crucial for visual effects (VFX), allowing transparent elements like smoke and reflections to blend seamlessly into scenes. We introduce TransPixeler, a method to extend pretrained video models for RGBA generation while retaining the original RGB capabilities. TransPixar leverages a diffusion transformer (DiT) architecture, incorporating alpha-specific tokens and using LoRA-based fine-tuning to jointly generate RGB and alpha channels with high consistency. By optimizing attention mechanisms, TransPixar preserves the strengths of the original RGB model and achieves strong alignment between RGB and alpha channels despite limited training data. Our approach effectively generates diverse and consistent RGBA videos, advancing the possibilities for VFX and interactive content creation.

文本生成视频透明度生成视觉特效扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。