从混音中生成两人面对面互动的3D面部动画,支持眼神交流和位置控制。
Talking Together: Synthesizing Co-Located 3D Conversations from Audio
- 双流架构+跨说话人注意力,分离混音并建模互动关系
- 引入眼动损失,实现自然互视;可文本控制相对位置与朝向
- 构建超200万对对话数据集,适用于虚拟现实与远程会话
我们解决从混合音频流中生成两个共处参与者完整3D面部动画的挑战。现有方法多生成类似视频会议的“悬浮说话头”,而本文首次显式建模真实面对面对话中关键的动态3D空间关系——包括相对位置、朝向及相互凝视。系统合成双方完整表演,包含精确口型同步,并可通过文本描述控制其相对头部姿态。为此提出双流架构,每一流负责一人输出,采用说话人角色嵌入与跨说话人交叉注意力机制以解耦混合音频并建模交互。此外,引入新型眼动损失以促进自然互视。为支撑数据密集型方法,构建新流水线,从真实场景视频中整理出超过200万对对话数据。所提方法生成流畅、可控且具空间感知的双人动画,显著优于现有基线,在感知真实感与交互连贯性上表现更优。
原文摘要 · Abstract (English)
We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often produce disembodied "talking heads" akin to a video conference call, our work is the first to explicitly model the dynamic 3D spatial relationship -- including relative position, orientation, and mutual gaze -- that is crucial for realistic in-person dialogues. Our system synthesizes the full performance of both individuals, including precise lip-sync, and uniquely allows their relative head poses to be controlled via textual descriptions. To achieve this, we propose a dual-stream architecture where each stream is responsible for one participant's output. We employ speaker's role embeddings and inter-speaker cross-attention mechanisms designed to disentangle the mixed audio and model the interaction. Furthermore, we introduce a novel eye gaze loss to promote natural, mutual eye contact. To power our data-hungry approach, we introduce a novel pipeline to curate a large-scale conversational dataset consisting of over 2 million dyadic pairs from in-the-wild videos. Our method generates fluid, controllable, and spatially aware dyadic animations suitable for immersive applications in VR and telepresence, significantly outperforming existing baselines in perceived realism and interaction coherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。