无需训练即可生成多角色语音驱动的高保真连贯视频
Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback
- 用扩散变换器+奖励反馈实现无训练多角色动画
- 支持三名以上角色语音同步,长视频生成更连贯
- 不依赖特殊数据集,可低成本部署于任意基础模型
近期扩散模型在语音驱动人物视频生成方面取得显著进展,质量与可控性超越传统方法。然而,现有方法仍面临唇形同步精度不足、长视频时间一致性差及多角色动画困难等问题。本文提出一种基于扩散变换器(DiT)的框架,实现任意长度逼真说话视频生成,并引入无训练的多角色语音驱动动画方法。首先,采用基于LoRA的训练策略结合位置偏移推理,实现高效长视频生成,同时保留基础模型能力;其次,通过部分参数更新与奖励反馈机制,提升唇形同步与自然肢体动作表现;最后,提出无训练的掩码分类引导(Mask-CFG)方法,无需特定数据集或模型修改,支持三名及以上角色的语音驱动动画。实验表明,该方法在高质量、时序一致性和多角色生成方面均优于现有最先进方法,以简单、高效、低成本方式实现目标。
原文摘要 · Abstract (English)
Recent advances in diffusion models have significantly improved audio-driven human video generation, surpassing traditional methods in both quality and controllability. However, existing approaches still face challenges in lip-sync accuracy, temporal coherence for long video generation, and multi-character animation. In this work, we propose a diffusion transformer (DiT)-based framework for generating lifelike talking videos of arbitrary length, and introduce a training-free method for multi-character audio-driven animation. First, we employ a LoRA-based training strategy combined with a position shift inference approach, which enables efficient long video generation while preserving the capabilities of the foundation model. Moreover, we combine partial parameter updates with reward feedback to enhance both lip synchronization and natural body motion. Finally, we propose a training-free approach, Mask Classifier-Free Guidance (Mask-CFG), for multi-character animation, which requires no specialized datasets or model modifications and supports audio-driven animation for three or more characters. Experimental results demonstrate that our method outperforms existing state-of-the-art approaches, achieving high-quality, temporally coherent, and multi-character audio-driven video generation in a simple, efficient, and cost-effective manner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。