arXiv:2507.01390cs.CV2025-07ICCV被引 20

解决极端情况下的说话头生成身份泄露与伪影问题

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases

  • 用增强运动指示器分离动作特征中的身份信息
  • 利用泄漏的身份信息修复渲染伪影,提升画质
  • 适合需要高保真说话头生成的场景或研究者

说话头生成在多个领域日益重要,但现有方法在极端情况下常出现身份泄露(IL)和渲染伪影(RA)。通过分析发现:(1)身份信息嵌入于动作特征中导致IL;(2)该信息可被用于缓解RA。为此提出FixTalk框架:引入增强运动指示器(EMI)有效解耦身份与动作特征,减轻身份泄露;设计增强细节指示器(EDI),利用泄漏的身份信息补充缺失细节,修复伪影。大量实验表明,FixTalk显著降低IL与RA,性能优于当前最优方法。

原文摘要 · Abstract (English)

Talking head generation is gaining significant importance across various domains, with a growing demand for high-quality rendering. However, existing methods often suffer from identity leakage (IL) and rendering artifacts (RA), particularly in extreme cases. Through an in-depth analysis of previous approaches, we identify two key insights: (1) IL arises from identity information embedded within motion features, and (2) this identity information can be leveraged to address RA. Building on these findings, this paper introduces FixTalk, a novel framework designed to simultaneously resolve both issues for high-quality talking head generation. Firstly, we propose an Enhanced Motion Indicator (EMI) to effectively decouple identity information from motion features, mitigating the impact of IL on generated talking heads. To address RA, we introduce an Enhanced Detail Indicator (EDI), which utilizes the leaked identity information to supplement missing details, thus fixing the artifacts. Extensive experiments demonstrate that FixTalk effectively mitigates IL and RA, achieving superior performance compared to state-of-the-art methods.

说话头生成身份泄露图像修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。