让3D虚拟人对话更自然,能无缝切换说与听。
DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations
- 统一框架同步生成说话与倾听的动态行为
- 在50小时多轮对话中实现流畅角色切换
- 适合需要真实交互感的虚拟人应用
面对面交流中,人们需自然切换说与听的角色。现有3D说话头生成模型仅关注单向表达,忽视互动动态,导致交互生硬、转换突兀。为此,我们提出新任务——多轮双说话人互动的3D说话头生成,要求模型同时处理并生成说与听行为。为此,我们提出DualTalk框架,统一建模说话者与倾听者的动态行为,不仅生成逼真的说话动作,还能持续输出生动的非语言反馈,有效捕捉角色间的互动关系。我们还构建了一个包含50小时多轮对话、超过1000个角色的新数据集,参与者持续切换说与听。大量实验表明,该方法显著提升了双说话人对话中3D说话头的自然度与表现力。建议观看补充视频:https://ziqiaopeng.github.io/dualtalk。
原文摘要 · Abstract (English)
In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive conversation, which leads to unnatural interactions and awkward transitions. To address this issue, we propose a new task -- multi-round dual-speaker interaction for 3D talking head generation -- which requires models to handle and generate both speaking and listening behaviors in continuous conversation. To solve this task, we introduce DualTalk, a novel unified framework that integrates the dynamic behaviors of speakers and listeners to simulate realistic and coherent dialogue interactions. This framework not only synthesizes lifelike talking heads when speaking but also generates continuous and vivid non-verbal feedback when listening, effectively capturing the interplay between the roles. We also create a new dataset featuring 50 hours of multi-round conversations with over 1,000 characters, where participants continuously switch between speaking and listening roles. Extensive experiments demonstrate that our method significantly enhances the naturalness and expressiveness of 3D talking heads in dual-speaker conversations. We recommend watching the supplementary video: https://ziqiaopeng.github.io/dualtalk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。