用多信号控制生成自然说话头视频,支持语音与表情同步
Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation
- 设计并行Mamba结构,各分支用不同信号控制面部不同区域
- 通过门控机制和掩码丢弃策略实现多信号无冲突驱动
- 在多个数据集上生成视频自然流畅,适配虚拟主播与交互系统
说话头合成对虚拟形象和人机交互至关重要。然而,现有方法通常仅支持单一模态控制,限制了实际应用。为此,我们提出端到端的视频扩散框架ACTalker,支持多信号与单信号控制的说话头视频生成。针对多信号控制,设计具有多分支的并行Mamba结构,每个分支使用独立驱动信号控制特定面部区域,并引入门控机制实现灵活控制。为确保时序与空间上的自然协调,采用Mamba结构使驱动信号可同时操控特征令牌的时间与空间维度。此外,提出掩码丢弃策略,使各驱动信号在对应面部区域独立控制,避免控制冲突。实验表明,该方法能生成由多种信号驱动的自然外观面部视频,且Mamba层可无缝融合多模态驱动信号而无冲突。项目主页见 https://harlanhong.github.io/publications/actalker/index.html。
原文摘要 · Abstract (English)
Talking head synthesis is vital for virtual avatars and human-computer interaction. However, most existing methods are typically limited to accepting control from a single primary modality, restricting their practical utility. To this end, we introduce \textbf{ACTalker}, an end-to-end video diffusion framework that supports both multi-signals control and single-signal control for talking head video generation. For multiple control, we design a parallel mamba structure with multiple branches, each utilizing a separate driving signal to control specific facial regions. A gate mechanism is applied across all branches, providing flexible control over video generation. To ensure natural coordination of the controlled video both temporally and spatially, we employ the mamba structure, which enables driving signals to manipulate feature tokens across both dimensions in each branch. Additionally, we introduce a mask-drop strategy that allows each driving signal to independently control its corresponding facial region within the mamba structure, preventing control conflicts. Experimental results demonstrate that our method produces natural-looking facial videos driven by diverse signals and that the mamba layer seamlessly integrates multiple driving modalities without conflict. The project website can be found at https://harlanhong.github.io/publications/actalker/index.html.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。