arXiv:2607.21434cs.CVcs.AI2026-07

通过自适应锚点闭环机制,提升视频换脸的稳定性与真实感。

Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervision in Video Face Swapping

  • 动态选择最差帧插入身份锚点,闭环优化生成质量。
  • 在相同预算下,自适应放置比均匀放置显著减少身份漂移。
  • 适合追求高保真换脸效果的研究者与工业应用开发者。

视频换脸缺乏自然的成对监督:现实中不存在某人以他人视频动作为基础的真实影像。当前最强方法 DreamID-V 的 SyncID-Pipe 通过替换真实片段中首尾两帧的身份,并仅用姿态序列重生成其余部分来构造配对。由于姿态不携带换入身份的外观信息,长视频、遮挡或极端姿态下,合成身份会长时间无锚定漂移;现有研究未分析锚点数量与位置的影响。本文提出自适应身份锚定(AIA):(i) 将合成器泛化至任意锚点集,适用于扩散强迫型变换器架构,条件化帧即固定其标记为零噪声;(ii) 通过闭环反馈机制评估每帧生成结果与真实参考身份的差异,在得分最低帧插入图像换脸锚点,直至通过阈值或耗尽预算;(iii) 复用该反馈作为自动数据过滤器。另一病理性问题——过度平滑皮肤的审美滤镜,根源同样在于微纹理未被目标函数所优化。因此,我们结合真实参照纹理恢复:从每帧非人脸区域匹配再着色,将真实视频中的子身份微纹理分频段迁移,并引入由视频自身频谱参考的第二通道进行谱域验证。我们认为身份锚密度是可调控的质量参数,并提出可验证实验:漂移-间隙曲线、相同预算下均匀与自适应放置对比、学生模型在 AIA 生成数据上的训练,以及含人类审美滤镜研究的纹理消融实验。

原文摘要 · Abstract (English)

Video face swapping has no natural paired supervision: no real footage exists of one person's face performing another person's video. The strongest current answer, DreamID-V's SyncID-Pipe, mints pairs by replacing the identity in exactly two frames of a real clip -- the first and the last -- and regenerating the rest from a pose sequence alone. Pose carries no appearance evidence of the swapped-in identity, so over long clips, occlusions, and extreme pose excursions the synthesized identity has a long unanchored span on which to drift; no published ablation examines anchor count or placement. We propose Adaptive Identity Anchoring (AIA): (i) generalize the synthesizer to arbitrary anchor sets, architecturally natural for diffusion-forcing-style transformers where conditioning on a frame is clamping its tokens to zero noise; (ii) place anchors by a closed feedback loop that scores every generated frame against the real reference identity and inserts an image-face-swapped anchor at the worst-scoring frame until the pair passes a threshold or exhausts a budget; (iii) reuse the loop's verdict as an automatic data filter. A second pathology, the beauty-filter look of over-smoothed skin, has the same root cause: micro-texture, like identity, is priced by none of the pipeline's objectives. We therefore pair AIA with Reality-Referenced Texture Restoration: matched re-graining from each real frame's non-face regions, band-split transfer of sub-identity micro-texture from the real footage, and a second, spectral acceptance channel refereed by the footage's own spectrum. Identity-anchor density, we argue, is a controllable quality dial, and we specify falsifiable experiments -- drift-versus-gap curves, uniform-versus-adaptive placement at matched budgets, student training on AIA-minted data, and texture ablations with a human beauty-filter study -- that would validate or refute the proposal.

视频换脸身份锚定闭环优化微纹理恢复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。