用文字和声音生成无限时长的高保真人物说话视频,支持多人多风格。
MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice
- 基于扩散Transformer架构,结合滑动窗口去噪策略实现无限视频生成。
- 10秒生成540x540p 10秒视频,720x720p仅需30秒,效率提升20倍。
- 支持多角色、多视角、口型同步与身份保持,适合影视动画生成场景。
我们提出MagicInfinite,一种新型扩散Transformer(DiT)框架,突破传统人物动画局限,可在真实人类、全身形象及动漫风格等多种角色类型上生成高质量视频。支持多种面部姿态,包括背面视角,并可通过输入掩码在多角色场景中精确指定发言人。通过三项创新解决核心挑战:(1) 3D全注意力机制结合滑动窗口去噪策略,实现跨多样角色风格的无限视频生成,保证时间连贯性与视觉质量;(2) 两阶段课程学习方案,融合音频实现口型同步、文本驱动表情动态、参考图像保持身份特征,实现长序列灵活多模态控制;(3) 区域特定掩码与自适应损失函数,平衡全局文本控制与局部音频引导,支持发言者专属动画。通过统一步数与配置缩放蒸馏技术提升效率,在8张H100 GPU上实现20倍推理加速:10秒内生成10秒540x540p视频或30秒生成720x720p视频,无质量损失。在新基准测试中,魔幻无限在音频-口型同步、身份保持与动作自然度方面均表现优异。项目已公开,地址为https://www.hedra.com/,示例见https://magicinfinite.github.io/。
原文摘要 · Abstract (English)
We present MagicInfinite, a novel diffusion Transformer (DiT) framework that overcomes traditional portrait animation limitations, delivering high-fidelity results across diverse character types-realistic humans, full-body figures, and stylized anime characters. It supports varied facial poses, including back-facing views, and animates single or multiple characters with input masks for precise speaker designation in multi-character scenes. Our approach tackles key challenges with three innovations: (1) 3D full-attention mechanisms with a sliding window denoising strategy, enabling infinite video generation with temporal coherence and visual quality across diverse character styles; (2) a two-stage curriculum learning scheme, integrating audio for lip sync, text for expressive dynamics, and reference images for identity preservation, enabling flexible multi-modal control over long sequences; and (3) region-specific masks with adaptive loss functions to balance global textual control and local audio guidance, supporting speaker-specific animations. Efficiency is enhanced via our innovative unified step and cfg distillation techniques, achieving a 20x inference speed boost over the basemodel: generating a 10 second 540x540p video in 10 seconds or 720x720p in 30 seconds on 8 H100 GPUs, without quality loss. Evaluations on our new benchmark demonstrate MagicInfinite's superiority in audio-lip synchronization, identity preservation, and motion naturalness across diverse scenarios. It is publicly available at https://www.hedra.com/, with examples at https://magicinfinite.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。