让多人对话视频生成更自然,解决音频与人物错配问题
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
- 用标签旋转位置编码实现音频与人物精准绑定
- 在多数据集上表现优于现有方法,支持指令跟随
- 适合需要多人交互视频生成的研究与应用
基于音频驱动的人体动画方法在生成同步面部动作和高质量视觉效果方面取得了显著进展。然而,现有方法主要针对单人动画,难以处理多路音频输入,存在音频与人物错配的问题,且在指令遵循能力上受限。为此,本文提出新任务:多人对话视频生成,并引入新框架 MultiTalk 来应对多人生成中的挑战。针对音频注入,研究多种方案并提出标签旋转位置编码(L-RoPE)方法,有效解决音频与人物绑定问题。训练过程中发现,部分参数训练和多任务学习对保持基础模型的指令遵循能力至关重要。MultiTalk 在多个数据集(包括说话头、说话身体和多人数据集)上均取得优异性能,验证了该方法强大的生成能力。
原文摘要 · Abstract (English)
Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing methods primarily focus on single human animation and struggle with multi-stream audio inputs, facing incorrect binding problems between audio and persons. Additionally, they exhibit limitations in instruction-following capabilities. To solve this problem, in this paper, we propose a novel task: Multi-Person Conversational Video Generation, and introduce a new framework, MultiTalk, to address the challenges during multi-person generation. Specifically, for audio injection, we investigate several schemes and propose the Label Rotary Position Embedding (L-RoPE) method to resolve the audio and person binding problem. Furthermore, during training, we observe that partial parameter training and multi-task training are crucial for preserving the instruction-following ability of the base model. MultiTalk achieves superior performance compared to other methods on several datasets, including talking head, talking body, and multi-person datasets, demonstrating the powerful generation capabilities of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。