提出评估说话人脸生成中身份泄漏的新方法,揭示参考图对唇动的影响。
Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework
- 设计三种测试场景检测唇部泄露:静音生成、音视频不匹配、音视频匹配
- 引入唇同步误差与静音音频唇动评分等新指标,量化泄漏程度
- 发现参考图像选择影响泄露,为模型设计提供实证依据
基于视频编辑的说话人脸生成旨在保持姿态、光照和手势等视频细节的同时,仅改变唇部动作,通常使用身份参考图像以维持说话人一致性。然而,该机制可能引发唇部泄漏,即生成的唇部动作受到参考图像影响,而非仅由驱动音频决定。这种泄漏难以通过标准指标和常规测试设置发现。为此,我们提出一种系统性评估方法,用于分析和量化唇部泄漏。框架包含三种互补的测试设置:静音输入生成、音视频不匹配配对、音视频匹配合成,并引入唇同步偏差及基于静音音频的唇同步得分等衍生指标。同时,研究不同身份参考选择对泄漏的影响,为参考设计提供洞见。所提方法具有模型无关性,可为未来说话人脸生成研究建立更可靠的基准。
原文摘要 · Abstract (English)
Video editing-based talking face generation aims to preserve video details such as pose, lighting, and gestures while modifying only lip motion, often using an identity reference image to maintain speaker consistency. However, this mechanism can introduce lip leakage, where generated lips are influenced by the reference image rather than solely by the driving audio. Such leakage is difficult to detect with standard metrics and conventional test setup. To address this, we propose a systematic evaluation methodology to analyze and quantify lip leakage. Our framework employs three complementary test setups: silent-input generation, mismatched audio-video pairing, and matched audio-video synthesis. We also introduce derived metrics including lip-sync discrepancy and silent-audio-based lip-sync scores. In addition, we study how different identity reference selections affect leakage, providing insights into reference design. The proposed methodology is model-agnostic and establishes a more reliable benchmark for future research in talking face generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。