用动态多模态检索实现低延迟个性化语音驱动人脸动画
Personalizing Causal Audio-Driven Facial Motion via Dynamic Multi-modal Retrieval

- 通过时序分层表示捕捉全局时序与高频细节,保持因果性
- 联合查询音频与动作动态提取风格先验,支持灵活模板数量
- 适用于需要实时互动的虚拟角色生成,如直播或元宇宙
语音驱动的人脸动画对沉浸式数字交互至关重要,但现有方法难以兼顾实时流处理与高保真个性化。当前技术常依赖引入延迟的音频前瞻,或要求用户预先编码静态嵌入,无法捕捉动态个性特征。本文提出一种端到端的因果框架,通过动态多模态风格检索实现个性化面部运动生成,支持超低延迟并充分利用非结构化风格参考。核心创新包括:(1) 时序分层运动表示,兼顾全局时序上下文与高频细节,同时维持解码因果性;(2) 多模态风格检索器,联合查询音频与运动以动态提取风格先验,不破坏因果性。该机制可支持可扩展的个性化,对模板数量与内容无限制。将上述组件集成至因果自回归架构中,本方法在唇音同步准确率、身份一致性及感知真实感方面显著优于现有最先进方法,经大量定量评估与用户研究验证。
原文摘要 · Abstract (English)
Audio-driven facial animation is essential for immersive digital interaction, yet existing frameworks fail to reconcile real-time streaming with high-fidelity personalization. Current methods often rely on latency-inducing audio look-ahead, or require high user compliance to pre-encode static embeddings that fails to capture dynamic idiosyncrasies. We present an end-to-end causal framework for personalizing causal facial motion generation via dynamic multi-modal style retrieval, enabling ultra-low latency while uniquely leveraging unstructured style references. We introduce two key innovations: (1) a temporal hierarchical motion representation that captures global temporal context and high-frequency details while maintaining decoding causality, and (2) a multi-modal style retriever that jointly queries audio and motion to dynamically extract stylistic priors without breaking causality. This mechanism allows for scalable personalization with total flexibility regarding the number and contents of templates. By integrating these components into a causal autoregressive architecture, our method significantly outperforms state-of-the-art approaches in lip-sync accuracy, identity consistency, and perceived realism, supported by extensive quantitative evaluations and user studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。