arXiv:2412.04037cs.CVcs.AI2024-12CVPR被引 42

让虚拟人物根据对话音频自动切换说与听的状态,实现自然互动。

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

  • 通过双人对话音频动态驱动头像在说话和倾听间切换
  • 在真实对话数据集上生成流畅、符合语境的面部动作
  • 适合需要自然人机交互的虚拟助手、社交机器人场景

设想与一个具备社交智能的代理进行对话,它能专注倾听你的言语,并及时提供视觉和语言反馈。这种无缝交互使多轮对话顺畅自然。为实现这一目标,我们提出 INFP——一种面向双人对话场景的音频驱动头像生成框架。不同于以往仅支持单向通信或需手动指定角色的模型,本模型可依据输入的双人音频,动态交替生成说话与倾听状态的头像。INFP 包含两个阶段:运动基头像模仿阶段,将真实对话视频中的面部交流行为映射到低维运动隐空间并驱动静态图像;音频引导运动生成阶段,通过去噪学习输入双人音频到运动隐码的映射,实现交互场景下的音频驱动头像生成。为推动该研究,我们构建了 DyConv——一个大规模互联网收集的丰富双人对话数据集。大量实验与可视化结果验证了方法的优越性能与有效性。

原文摘要 · Abstract (English)

Imagine having a conversation with a socially intelligent agent. It can attentively listen to your words and offer visual and linguistic feedback promptly. This seamless interaction allows for multiple rounds of conversation to flow smoothly and naturally. In pursuit of actualizing it, we propose INFP, a novel audio-driven head generation framework for dyadic interaction. Unlike previous head generation works that only focus on single-sided communication, or require manual role assignment and explicit role switching, our model drives the agent portrait dynamically alternates between speaking and listening state, guided by the input dyadic audio. Specifically, INFP comprises a Motion-Based Head Imitation stage and an Audio-Guided Motion Generation stage. The first stage learns to project facial communicative behaviors from real-life conversation videos into a low-dimensional motion latent space, and use the motion latent codes to animate a static image. The second stage learns the mapping from the input dyadic audio to motion latent codes through denoising, leading to the audio-driven head generation in interactive scenarios. To facilitate this line of research, we introduce DyConv, a large scale dataset of rich dyadic conversations collected from the Internet. Extensive experiments and visualizations demonstrate superior performance and effectiveness of our method. Project Page: https://grisoon.github.io/INFP/.

音频驱动头像生成双人对话交互系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。