多视角参考下保持大角度人脸一致性,避免生成僵硬复制问题。
Identity-Consistent Video Generation under Large Facial-Angle Variations

- 用多视角参考+区域掩码训练,避免模型偷懒学表面特征。
- 在大角度变化下身份一致率提升,动作自然度不下降。
- 适合需要高保真人脸视频生成的研究者或工业应用。
单视图参考生成视频的方法在大幅人脸角度变化下难以保持身份一致性。为解决此问题,本文提出无需交叉配对数据的多视角条件框架 $ ext{Mv}^2 ext{ID}$。通过区域掩码训练策略防止捷径学习,促使模型融合多视角互补的身份线索以提取本质特征;设计参考解耦-旋转位置编码机制,为视频与条件令牌分配不同位置编码,更好建模其异质性。此外,构建大规模包含多样化人脸角度变化的数据集,并提出针对身份一致性和动作自然性的专用评估指标。大量实验表明,该方法在保持动作自然度的同时显著提升身份一致性,优于使用交叉配对数据训练的现有方法。
原文摘要 · Abstract (English)
Single-view reference-to-video methods often struggle to preserve identity consistency under large facial-angle variations. This limitation naturally motivates the incorporation of multi-view facial references. However, simply introducing additional reference images exacerbates the \textit{copy-paste} problem, particularly the \textbf{\textit{view-dependent copy-paste}} artifact, which reduces facial motion naturalness. Although cross-paired data can alleviate this issue, collecting such data is costly. To balance the consistency and naturalness, we propose $\mathrm{Mv}^2\mathrm{ID}$, a multi-view conditioned framework under in-paired supervision. We introduce a region-masking training strategy to prevent shortcut learning and extract essential identity features by encouraging the model to aggregate complementary identity cues across views. In addition, we design a reference decoupled-RoPE mechanism that assigns distinct positional encoding to video and conditioning tokens for better modeling of their heterogeneous properties. Furthermore, we construct a large-scale dataset with diverse facial-angle variations and propose dedicated evaluation metrics for identity consistency and motion naturalness. Extensive experiments demonstrate that our method significantly improves identity consistency while maintaining motion naturalness, outperforming existing approaches trained with cross-paired data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。