arXiv:2504.04010cs.CVcs.LG2025-04ICCV被引 12

用扩散模型生成高保真、自然流畅的对话倾听者视频,支持长时连续输出。

DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion

  • 采用多模态条件扩散架构,融合语音与表情输入生成连贯头部动作。
  • 在RealTalk数据集上FID提升73.8%,VICO数据集运动表征准确率提高6.1%。
  • 适合需要高质量虚拟人物交互的影视、教育或元宇宙场景。

生成自然且细腻的长期互动中倾听者的动作仍是一个开放问题。现有方法通常依赖低维动作码生成面部行为,再进行逼真渲染,限制了视觉保真度和表现力。为此,我们提出DiTaiListener,基于多模态条件的视频扩散模型。首先,DiTaiListener-Gen根据说话人语音与面部动作生成短段倾听者响应;随后,DiTaiListener-Edit通过视频到视频扩散模型优化过渡帧,实现无缝衔接。具体地,DiTaiListener-Gen引入因果时序多模态适配器(CTM-Adapter),以因果方式整合说话人的听觉与视觉输入至扩散变换器(DiT)中,确保时间上一致的回应。针对长视频生成,我们设计了过渡优化模块,将多个短片段融合为连续流畅的视频,保持表情连贯性与图像质量。定量评估显示,DiTaiListener在基准数据集上达到当前最优表现:RealTalk上FID提升73.8%,VICO上FD指标提高6.1%。用户研究证实其在反馈、多样性与平滑性方面显著优于竞品。

原文摘要 · Abstract (English)

Generating naturalistic and nuanced listener motions for extended interactions remains an open problem. Existing methods often rely on low-dimensional motion codes for facial behavior generation followed by photorealistic rendering, limiting both visual fidelity and expressive richness. To address these challenges, we introduce DiTaiListener, powered by a video diffusion model with multimodal conditions. Our approach first generates short segments of listener responses conditioned on the speaker's speech and facial motions with DiTaiListener-Gen. It then refines the transitional frames via DiTaiListener-Edit for a seamless transition. Specifically, DiTaiListener-Gen adapts a Diffusion Transformer (DiT) for the task of listener head portrait generation by introducing a Causal Temporal Multimodal Adapter (CTM-Adapter) to process speakers' auditory and visual cues. CTM-Adapter integrates speakers' input in a causal manner into the video generation process to ensure temporally coherent listener responses. For long-form video generation, we introduce DiTaiListener-Edit, a transition refinement video-to-video diffusion model. The model fuses video segments into smooth and continuous videos, ensuring temporal consistency in facial expressions and image quality when merging short video segments produced by DiTaiListener-Gen. Quantitatively, DiTaiListener achieves the state-of-the-art performance on benchmark datasets in both photorealism (+73.8% in FID on RealTalk) and motion representation (+6.1% in FD metric on VICO) spaces. User studies confirm the superior performance of DiTaiListener, with the model being the clear preference in terms of feedback, diversity, and smoothness, outperforming competitors by a significant margin.

视频生成扩散模型多模态虚拟角色

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。