用关键点控制面部表情,实现高保真语音驱动动画与精准对口型。
KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization

- 基于关键点条件的流模型,分离嘴部动作与上半脸表情。
- 在数据有限条件下仍保持高精度对口型与风格一致性。
- 适合影视配音、虚拟人对话等需要精细控制的场景。
语音驱动的3D面部动画在追求高保真运动与精确艺术控制方面面临挑战。现有可控模型多依赖大规模、低质量的野外数据集学习全局风格控制,影响整体动画真实感。此外,这些框架常缺乏对话定位等任务所需的帧级时间精度,而匹配特定面部表情与唇音同步同样重要。我们提出KM-Speaker(关键点匹配说话者),一种新型的关键点条件流生成框架,从参考表演中提供全局风格引导与帧级时间控制。我们设计了一种解耦策略,将音频驱动的嘴部运动与关键点驱动的上半脸动态分离,并引入全局风格上下文保留机制,确保全脸表达的一致性。在数据受限环境下,KM-Speaker显著提升动画保真度与可控性,在唇音同步准确率、风格遵循度与表达时间控制方面持续优于当前最优方法。
原文摘要 · Abstract (English)
Speech-driven 3D facial animation methods face significant challenges in simultaneously achieving high-fidelity motion and precise artistic control at production quality. Existing controllable models typically learn global style control by relying on large-scale, low-quality \emph{in-the-wild} datasets that compromise overall animation realism. Furthermore, these frameworks often lack the fine-grained temporal precision required for demanding tasks such as dialogue localization (e.g., dubbing), where matching specific facial expressions is as critical as lip synchronization. We present KM-Speaker (Keypoint-Matching Speaker), a novel keypoint-conditioned flow-based generative framework that provides both global style guidance and frame-level temporal control from reference performances. We propose a disentanglement strategy that separates audio-driven lip motion from keypoint-driven upper-face dynamics, together with a global style context preservation mechanism to ensure coherent full-face expressiveness. KM-Speaker advances example-based 3D facial animation by achieving high-fidelity motion and flexible controllability in a data-constrained setting, consistently outperforming state-of-the-art methods in lip-sync accuracy, style adherence, and expressive temporal control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。