用多张人脸照片生成个性化动态视频,支持精准身份控制。
Ingredients: Blending Custom Photos with Video Diffusion Transformers
- 通过多尺度投影将人脸特征映射到视频扩散模型的上下文空间
- 动态路由机制实现多个身份嵌入在时空区域的智能分配
- 支持基于自定义照片生成高质量个性视频,适合内容创作者使用
本文提出一种名为Ingredients的框架,利用视频扩散Transformer实现多张特定身份照片的视频定制。方法包含三个核心模块:(i) 多视角人脸提取器,从全局与局部捕捉每个身份的精细面部特征;(ii) 多尺度投影器,将人脸嵌入映射至视频扩散模型的图像查询上下文空间;(iii) 身份路由器,动态分配多个身份嵌入至对应时空区域。依托精心构建的文本-视频数据集与多阶段训练协议,Ingredients在将自定义照片转为动态个性化视频方面表现优异。定性评估显示,该方法在基于Transformer的生成视频控制中具有显著优势,优于现有技术。代码、数据及模型权重已公开于https://github.com/feizc/Ingredients。
原文摘要 · Abstract (English)
This paper presents a powerful framework to customize video creations by incorporating multiple specific identity (ID) photos, with video diffusion Transformers, referred to as Ingredients. Generally, our method consists of three primary modules: (i) a facial extractor that captures versatile and precise facial features for each human ID from both global and local perspectives; (ii) a multi-scale projector that maps face embeddings into the contextual space of image query in video diffusion transformers; (iii) an ID router that dynamically combines and allocates multiple ID embedding to the corresponding space-time regions. Leveraging a meticulously curated text-video dataset and a multi-stage training protocol, Ingredients demonstrates superior performance in turning custom photos into dynamic and personalized video content. Qualitative evaluations highlight the advantages of proposed method, positioning it as a significant advancement toward more effective generative video control tools in Transformer-based architecture, compared to existing methods. The data, code, and model weights are publicly available at: https://github.com/feizc/Ingredients.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。