让虚拟头像实时互动,说话、点头、笑都能即时响应。
Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
- 用扩散强制机制实现低延迟实时反应,支持语音和动作输入。
- 无需额外标注数据,合成负样本训练出生动自然的表情动作。
- 交互延迟约500毫秒,速度比基线快6.8倍,用户偏好超80%。
说话头生成技术可将静态肖像转化为用于虚拟通信与内容创作的逼真虚拟形象。然而,现有模型仍难以呈现真正互动的体验,常产生单向回应且缺乏情感投入。我们识别出两大核心挑战:在因果约束下实现实时运动生成,以及在无额外标注数据的情况下学习富有表现力的互动反应。为此,我们提出Avatar Forcing框架,通过扩散强制建模实时用户-虚拟形象交互。该设计使虚拟形象能够处理实时多模态输入(如音频与动作),以低延迟对语言及非语言线索(如讲话、点头、笑声)做出即时响应。此外,我们引入直接偏好优化方法,利用移除用户条件构造的合成负样本,实现无需标签的表达性互动学习。实验表明,本框架实现低延迟实时交互(约500毫秒),较基线提速6.8倍,生成的动态更具反应性与表现力,在用户偏好测试中胜出超过80%。
原文摘要 · Abstract (English)
Talking head generation creates lifelike avatars from static portraits for virtual communication and content creation. However, current models do not yet convey the feeling of truly interactive communication, often generating one-way responses that lack emotional engagement. We identify two key challenges toward truly interactive avatars: generating motion in real-time under causal constraints and learning expressive, vibrant reactions without additional labeled data. To address these challenges, we propose Avatar Forcing, a new framework for interactive head avatar generation that models real-time user-avatar interactions through diffusion forcing. This design allows the avatar to process real-time multimodal inputs, including the user's audio and motion, with low latency for instant reactions to both verbal and non-verbal cues such as speech, nods, and laughter. Furthermore, we introduce a direct preference optimization method that leverages synthetic losing samples constructed by dropping user conditions, enabling label-free learning of expressive interaction. Experimental results demonstrate that our framework enables real-time interaction with low latency (approximately 500ms), achieving 6.8X speedup compared to the baseline, and produces reactive and expressive avatar motion, which is preferred over 80% against the baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。