让语音驱动的虚拟人脸说话更自然,支持风格与情绪精细调控。
PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
- 通过隐式关键点变形实现逐词级说话风格编辑
- 可调节唇部动作幅度和情感强度,生成多样化表情
- 适合需要个性化虚拟人表达的应用场景
近年来语音驱动的人脸动画生成在唇音同步方面取得显著进展,但现有方法对表情风格和情感表达的控制不足,导致输出结果单调。本文聚焦唇音对齐与情感控制两大核心问题,提出PC-Talk框架,通过隐式关键点形变实现精准面部动画调控。首先,唇音对齐模块可在词级别精确编辑说话风格,并调节唇部运动幅度以模拟不同音量,同时保持与音频的同步;其次,情感控制模块生成真实感十足的情感特征,支持情感强度的精细调节及多情绪在不同面部区域的组合。大量实验表明,该方法在HDTF与MEAD数据集上均达到当前最优性能,展现出卓越的可控性与多样性。
原文摘要 · Abstract (English)
Recent advancements in audio-driven talking face generation have made great progress in lip synchronization. However, current methods often lack sufficient control over facial animation such as speaking style and emotional expression, resulting in uniform outputs. In this paper, we focus on improving two key factors: lip-audio alignment and emotion control, to enhance the diversity and user-friendliness of talking videos. Lip-audio alignment control focuses on elements like speaking style and the scale of lip movements, whereas emotion control is centered on generating realistic emotional expressions, allowing for modifications in multiple attributes such as intensity. To achieve precise control of facial animation, we propose a novel framework, PC-Talk, which enables lip-audio alignment and emotion control through implicit keypoint deformations. First, our lip-audio alignment control module facilitates precise editing of speaking styles at the word level and adjusts lip movement scales to simulate varying vocal loudness levels, maintaining lip synchronization with the audio. Second, our emotion control module generates vivid emotional facial features with pure emotional deformation. This module also enables the fine modification of intensity and the combination of multiple emotions across different facial regions. Our method demonstrates outstanding control capabilities and achieves state-of-the-art performance on both HDTF and MEAD datasets in extensive experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。