让多个角色同时入镜,生成身份清晰且动作自然的视频。
GroupVideo: Multi-Identity Customized Text-to-Video Generation

- 用多张人脸图联合编码,实现多身份精准对齐。
- 生成20,000条高质量多角色视频,突破数据瓶颈。
- 适合需要多人视频生成的研究与应用开发者。
当前的身份定制化视频生成方法主要局限于单身份场景,因缺乏显式的身份分离机制,多身份设置下常出现身份混淆。现有方法通过拼接人脸图像作为输入条件,往往导致面部表情和动作不自然,呈现“复制粘贴”现象。为此,我们提出GroupVideo,一种基于视频扩散变换器的新框架,利用多张个体照片生成身份定制视频。该框架引入多模态身份对齐:视觉对齐联合编码多张人脸以提供鲁棒的身份参考,语义对齐则通过语义感知器提升动作自然度。此外,引入具有空间引导的ID定位模块,结合边界框约束和掩码正则化损失,聚焦面部区域,增强身份保真度并提高训练效率。针对多身份视频数据集稀缺问题,我们构建了一个包含20,000条视频的高质量数据集,为未来研究奠定基础。大量实验表明,GroupVideo在生成多角色视频时,身份一致性与动作自然性均优于现有方法。
原文摘要 · Abstract (English)
Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity separation mechanisms often leads to identity confusion in multi-identity settings. Existing multi-identity approaches, which directly extend single-identity frameworks by concatenating face images as input conditions, frequently result in unnatural facial expressions and motions, manifesting as the "copy-paste" phenomenon. To overcome these limitations, we introduce GroupVideo, a novel framework that leverages multiple individual photographs to generate identitycustomized video. Built upon Video Diffusion Transformers, GroupVideo incorporates multimodal identity alignment: visual alignment jointly encodes multiple face images to provide robust identity references, while semantic alignment introduces a semantic perceiver to enhance the naturalness of motions. An ID localization module with spatial guidance is introduced to address identity blending and enhance identity fidelity, along with bounding box constraints and mask regularization loss, to focus on facial regions and improve training efficiency. In response to the shortage of multi-ID video datasets, we have curated a comprehensive high-quality dataset of 20,000 videos, thereby establishing a crucial resource to advance future research in multi-ID video generation. Extensive experiments demonstrate that GroupVideo outperforms existing methods in generating multi-character videos with consistent identities and natural motions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。