用参考视频保持人物身份,生成自然动态的视频。
Slot-ID: Identity-Preserving Video Generation from Reference Videos via Slot-Based Temporal Identity Encoding
- 用短参考视频替代单张图片,捕捉动态特征。
- 在大幅姿态变化下仍能准确保留身份特征。
- 适合需要高保真人物形象的视频生成场景。
生成忠实于提示且保留用户指定身份的视频仍具挑战:模型需从稀疏参考中推断面部动态,同时平衡身份保持与运动自然性之间的矛盾。仅以单张图像为条件会忽略时间特征,导致动作僵硬、变形不自然及视角或表情变化时出现“平均化”人脸。为此,我们提出一种基于扩散-变压器架构的视频生成器,采用短参考视频而非单张肖像作为条件。核心思路是引入参考视频中的动态信息,一段短片段可揭示特定主体的表情形成模式(如微笑如何随姿态和光照变化)。通过Sinkhorn路由编码器学习紧凑的身份令牌,既捕捉特征动态又兼容预训练主干网络。尽管仅增加轻量级条件,该方法在大幅姿态变化和丰富表情表现下显著提升身份保留能力,同时保持提示忠实度与视觉真实性,适用于多样主体与提示。
原文摘要 · Abstract (English)
Producing prompt-faithful videos that preserve a user-specified identity remains challenging: models need to extrapolate facial dynamics from sparse reference while balancing the tension between identity preservation and motion naturalness. Conditioning on a single image completely ignores the temporal signature, which leads to pose-locked motions, unnatural warping, and "average" faces when viewpoints and expressions change. To this end, we introduce an identity-conditioned variant of a diffusion-transformer video generator which uses a short reference video rather than a single portrait. Our key idea is to incorporate the dynamics in the reference. A short clip reveals subject-specific patterns, e.g., how smiles form, across poses and lighting. From this clip, a Sinkhorn-routed encoder learns compact identity tokens that capture characteristic dynamics while remaining pretrained backbone-compatible. Despite adding only lightweight conditioning, the approach consistently improves identity retention under large pose changes and expressive facial behavior, while maintaining prompt faithfulness and visual realism across diverse subjects and prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。