arXiv:2411.19509cs.CVcs.LG2024-11被引 44

Ditto实现可控且实时的虚拟人说话头生成。

Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis

  • 在特定运动空间中用扩散模型生成面部动作表示,提升控制精度。
  • 支持实时推理与低首帧延迟,满足交互应用需求。
  • 可灵活调节表情与动作,适合智能助手等实时场景。

近期扩散模型提升了说话头合成的表情细腻度和头部动作生动性,但也导致推理速度慢、生成结果难以控制。为此,我们提出Ditto,一种基于扩散模型的说话头框架,实现细粒度控制与实时推理。具体而言,采用现成的运动提取器,并设计扩散变压器,在特定运动空间中生成表示。通过优化模型架构与训练策略,解决运动与身份未充分解耦、表示内部差异大的问题。同时,引入多样条件信号,建立运动表示与面部语义间的映射,实现生成过程调控与结果修正。此外,联合优化整体框架,支持流式处理、实时推理与低首帧延迟,适用于智能助手等交互应用。大量实验表明,Ditto生成的说话头视频表现优异,在可控性与实时性能上均具优势。

原文摘要 · Abstract (English)

Recent advances in diffusion models have endowed talking head synthesis with subtle expressions and vivid head movements, but have also led to slow inference speed and insufficient control over generated results. To address these issues, we propose Ditto, a diffusion-based talking head framework that enables fine-grained controls and real-time inference. Specifically, we utilize an off-the-shelf motion extractor and devise a diffusion transformer to generate representations in a specific motion space. We optimize the model architecture and training strategy to address the issues in generating motion representations, including insufficient disentanglement between motion and identity, and large internal discrepancies within the representation. Besides, we employ diverse conditional signals while establishing a mapping between motion representation and facial semantics, enabling control over the generation process and correction of the results. Moreover, we jointly optimize the holistic framework to enable streaming processing, real-time inference, and low first-frame delay, offering functionalities crucial for interactive applications such as AI assistants. Extensive experimental results demonstrate that Ditto generates compelling talking head videos and exhibits superiority in both controllability and real-time performance.

说话头生成扩散模型实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。