用注意力机制实现高保真个性化配音,保留说话人特征。
PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
- 分两阶段:先建模说话风格驱动几何,再用双注意力渲染面部纹理。
- 在LRS3数据集上唇动同步准确率达92.1%,视觉质量优于现有方法。
- 无需特定人训练,通用框架也能媲美专用模型,适合跨人物配音场景。
针对音频驱动的视觉配音任务,如何在保证精确唇动同步的同时保持说话人独特风格仍是挑战。现有方法难以捕捉说话人的个性表达或保留面部细节。本文提出PersonaTalk,一种基于注意力的两阶段框架,包括几何构建与面部渲染。第一阶段设计风格感知的音频编码模块,通过交叉注意力将说话风格注入音频特征,并驱动说话人模板几何生成唇动同步的几何形态。第二阶段引入双注意力面部渲染器,包含唇部注意力和面部注意力两个并行交叉注意力层,分别从参考帧中采样纹理以合成完整面部。该设计有效保留了精细面部特征。大量实验与用户研究证明,PersonaTalk在视觉质量、唇动同步精度和说话人特征保留方面均优于当前最优方法。此外,作为通用框架,其性能可媲美专门针对个体训练的先进方法。
原文摘要 · Abstract (English)
For audio-driven visual dubbing, it remains a considerable challenge to uphold and highlight speaker's persona while synthesizing accurate lip synchronization. Existing methods fall short of capturing speaker's unique speaking style or preserving facial details. In this paper, we present PersonaTalk, an attention-based two-stage framework, including geometry construction and face rendering, for high-fidelity and personalized visual dubbing. In the first stage, we propose a style-aware audio encoding module that injects speaking style into audio features through a cross-attention layer. The stylized audio features are then used to drive speaker's template geometry to obtain lip-synced geometries. In the second stage, a dual-attention face renderer is introduced to render textures for the target geometries. It consists of two parallel cross-attention layers, namely Lip-Attention and Face-Attention, which respectively sample textures from different reference frames to render the entire face. With our innovative design, intricate facial details can be well preserved. Comprehensive experiments and user studies demonstrate our advantages over other state-of-the-art methods in terms of visual quality, lip-sync accuracy and persona preservation. Furthermore, as a person-generic framework, PersonaTalk can achieve competitive performance as state-of-the-art person-specific methods. Project Page: https://grisoon.github.io/PersonaTalk/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。