arXiv:2502.13995cs.GRcs.CV2025-02被引 22

无需微调,用3D人脸结构增强生成视频面部动态与身份一致性。

FantasyID: Face Knowledge Enhanced ID-Preserving Video Generation

  • 引入3D人脸几何先验保证生成面部结构合理。
  • 多视角人脸增强提升表情与姿态多样性,避免复制粘贴式生成。
  • 自适应融合机制按层选择性注入特征,平衡身份与动态建模。

无需微调的预训练视频扩散模型在保持身份一致的文本到视频生成(IPT2V)中因高效与可扩展性而受到关注。然而,在保持身份不变的同时实现满意的面部动态仍面临挑战。本文提出一种新型无微调IPT2V框架FantasyID,通过在扩散变换器(DiT)基础上增强人脸知识。核心是引入3D人脸几何先验,确保视频合成中面部结构合理。为防止模型采用简单复制参考人脸的捷径,设计多视角人脸增强策略,捕捉多样化的2D面部外观特征,从而提升表情与头部姿态的动态表现。此外,融合2D与3D特征后,不使用简单的交叉注意力注入,而是采用可学习的分层自适应机制,有选择地将融合特征注入各层DiT,促进身份保留与运动动态的均衡建模。实验结果验证了本模型在当前无微调IPT2V方法中的优越性。

原文摘要 · Abstract (English)

Tuning-free approaches adapting large-scale pre-trained video diffusion models for identity-preserving text-to-video generation (IPT2V) have gained popularity recently due to their efficacy and scalability. However, significant challenges remain to achieve satisfied facial dynamics while keeping the identity unchanged. In this work, we present a novel tuning-free IPT2V framework by enhancing face knowledge of the pre-trained video model built on diffusion transformers (DiT), dubbed FantasyID. Essentially, 3D facial geometry prior is incorporated to ensure plausible facial structures during video synthesis. To prevent the model from learning copy-paste shortcuts that simply replicate reference face across frames, a multi-view face augmentation strategy is devised to capture diverse 2D facial appearance features, hence increasing the dynamics over the facial expressions and head poses. Additionally, after blending the 2D and 3D features as guidance, instead of naively employing cross-attention to inject guidance cues into DiT layers, a learnable layer-aware adaptive mechanism is employed to selectively inject the fused features into each individual DiT layers, facilitating balanced modeling of identity preservation and motion dynamics. Experimental results validate our model's superiority over the current tuning-free IPT2V methods.

视频生成身份保持扩散模型3D先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。