让文字动画同时保持清晰外观与透明效果,无需重训练模型。
TransText: Alpha-as-RGB Representation for Transparent Text Animation
- 将透明度编码为兼容RGB的视觉信号,避免修改预训练模型。
- 生成高质量透明文字动画,支持细腻动态效果。
- 适合需要精细文字动画设计的视觉创作场景。
我们提出首个将图像到视频模型适配至分层文本(字形)动画的方法,该能力对动态视觉设计至关重要。现有方法通常将透明度(alpha通道)作为附加于RGB空间的额外潜在维度处理,需重建基于RGB的变分自编码器(VAE),但因高质量透明字形数据稀缺,重训练成本高且可能破坏从大规模RGB语料中学到的语义先验,导致潜在特征混杂。为此,我们提出TransText框架,基于创新的Alpha-as-RGB范式,在不修改预训练生成流形的前提下联合建模外观与透明度。TransText通过潜在空间拼接将alpha通道嵌入为兼容RGB的视觉信号,显式保证RGB与α之间的跨模态一致性,防止特征纠缠。实验表明,TransText显著优于基线方法,能生成连贯、高保真且具有多样化细微效果的透明文字动画。
原文摘要 · Abstract (English)
We introduce the first method, to the best of our knowledge, for adapting image-to-video models to layer-aware text (glyph) animation, a capability critical for practical dynamic visual design. Existing approaches predominantly handle the transparency-encoding (alpha channel) as an extra latent dimension appended to the RGB space, necessitating the reconstruction of the underlying RGB-centric variational autoencoder (VAE). However, given the scarcity of high-quality transparent glyph data, retraining the VAE is computationally expensive and may erode the robust semantic priors learned from massive RGB corpora, potentially leading to latent pattern mixing. To mitigate these limitations, we propose TransText, a framework based on a novel Alpha-as-RGB paradigm to jointly model appearance and transparency without modifying the pre-trained generative manifold. TransText embeds the alpha channel as an RGB-compatible visual signal through latent spatial concatenation, explicitly ensuring strict cross-modal (RGB-and-Alpha) consistency while preventing feature entanglement. Our experiments demonstrate that TransText significantly outperforms baselines, generating coherent, high-fidelity transparent animations with diverse, fine-grained effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。