arXiv:2506.19833cs.CV2025-06被引 6

让多个角色在同一场景中自然对话,通过动态3D掩码控制谁在说话。

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router

  • 用细粒度嵌入路由机制绑定说话人与语音,解决多人对话的对应问题。
  • 提出3D掩码嵌入路由,实现帧级精细控制,提升掩码准确性和时序平滑性。
  • 首个专为多角色对话视频设计的数据集,适合研究多主体交互生成。

近年来,音频驱动的说话头生成取得了显著进展,但现有方法主要集中在单角色场景。尽管部分方法可生成两人间的独立对话视频,但如何在相同空间环境中生成多个真实共现角色的统一对话视频仍缺乏有效解决方案。该任务面临两大挑战:音频到角色的对应控制,以及缺少包含同一场景下多角色对话的合适数据集。为此,我们提出Bind-Your-Avatar,一种基于MM-DiT的模型,专用于同一场景下的多角色对话视频生成。具体包括:(1) 提出细粒度嵌入路由框架,将‘谁’与‘说什么’绑定,解决音频-角色对应控制问题;(2) 设计两种3D掩码嵌入路由方法,实现帧级精细化控制,结合几何先验的损失函数和掩码优化策略,提升预测掩码的准确性与时序平滑性;(3) 构建首个面向多角色对话视频生成的公开数据集,并提供开源数据处理流程;(4) 建立双角色对话视频生成基准,实验表明其性能优于多个前沿方法。

原文摘要 · Abstract (English)

Recent years have witnessed remarkable advances in audio-driven talking head generation. However, existing approaches predominantly focus on single-character scenarios. While some methods can create separate conversation videos between two individuals, the critical challenge of generating unified conversation videos with multiple physically co-present characters sharing the same spatial environment remains largely unaddressed. This setting presents two key challenges: audio-to-character correspondence control and the lack of suitable datasets featuring multi-character talking videos within the same scene. To address these challenges, we introduce Bind-Your-Avatar, an MM-DiT-based model specifically designed for multi-talking-character video generation in the same scene. Specifically, we propose (1) A novel framework incorporating a fine-grained Embedding Router that binds `who' and `speak what' together to address the audio-to-character correspondence control. (2) Two methods for implementing a 3D-mask embedding router that enables frame-wise, fine-grained control of individual characters, with distinct loss functions based on observed geometric priors and a mask refinement strategy to enhance the accuracy and temporal smoothness of the predicted masks. (3) The first dataset, to the best of our knowledge, specifically constructed for multi-talking-character video generation, and accompanied by an open-source data processing pipeline, and (4) A benchmark for the dual-talking-characters video generation, with extensive experiments demonstrating superior performance over multiple state-of-the-art methods.

多角色生成3D掩码对话视频音频驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。