无需额外参数,实现多身份视频生成与编辑的统一框架。
Concat-ID: Towards Universal Identity-Preserving Video Synthesis
- 用变分自编码器提取图像特征,拼接至视频隐变量序列中。
- 在多身份场景下保持身份一致性,生成视频自然度更高。
- 适合虚拟试穿、背景可控等实际应用,可扩展性强。
我们提出 Concat-ID,一种统一的身份保真视频生成框架。该方法利用变分自编码器提取图像特征,并将其沿序列维度拼接至视频隐变量中,仅依赖固有的3D自注意力机制实现融合,无需额外参数或模块。引入新颖的跨视频配对策略和多阶段训练方案,在保证身份一致性的同时提升面部可编辑性与视频自然度。大量实验表明,Concat-ID在单身份与多身份生成任务中均优于现有方法,且可无缝扩展至多主体场景,包括虚拟试穿与背景可控生成。该方法为身份保真视频合成建立了新基准,提供了一种通用且可扩展的解决方案。
原文摘要 · Abstract (English)
We present Concat-ID, a unified framework for identity-preserving video generation. Concat-ID employs variational autoencoders to extract image features, which are then concatenated with video latents along the sequence dimension. It relies exclusively on inherent 3D self-attention mechanisms to incorporate them, eliminating the need for additional parameters or modules. A novel cross-video pairing strategy and a multi-stage training regimen are introduced to balance identity consistency and facial editability while enhancing video naturalness. Extensive experiments demonstrate Concat-ID's superiority over existing methods in both single and multi-identity generation, as well as its seamless scalability to multi-subject scenarios, including virtual try-on and background-controllable generation. Concat-ID establishes a new benchmark for identity-preserving video synthesis, providing a versatile and scalable solution for a wide range of applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。