用结构化扩散模型生成多人互动的高质视频
Multi-identity Human Image Animation with Structural Video Diffusion
- 通过身份嵌入保持多人外观一致性
- 融合深度与法向图建模人物与物体交互
- 适用于复杂多人场景视频生成任务
从单张图像生成高质量、可控的人类视频是一项挑战,尤其在涉及多人及与物体互动的复杂场景中。现有方法虽在单人场景有效,却难以处理多身份交互的关联性与3D动态分布建模。本文提出「Structural Video Diffusion」框架,引入身份专属嵌入以保持个体外观一致性,并设计结构学习机制,结合深度图与表面法向信息建模人-物交互。同时,扩充现有数据集,新增25,000个包含多样化多人与物体交互的视频样本,为训练提供坚实基础。实验表明,该方法在生成多主体动态、丰富的真实感视频方面表现优异,显著提升人类中心视频生成水平。代码已开源。
原文摘要 · Abstract (English)
Generating human videos from a single image while ensuring high visual quality and precise control is a challenging task, especially in complex scenarios involving multiple individuals and interactions with objects. Existing methods, while effective for single-human cases, often fail to handle the intricacies of multi-identity interactions because they struggle to associate the correct pairs of human appearance and pose condition and model the distribution of 3D-aware dynamics. To address these limitations, we present \emph{Structural Video Diffusion}, a novel framework designed for generating realistic multi-human videos. Our approach introduces two core innovations: identity-specific embeddings to maintain consistent appearances across individuals and a structural learning mechanism that incorporates depth and surface-normal cues to model human-object interactions. Additionally, we expand existing human video dataset with 25K new videos featuring diverse multi-human and object interaction scenarios, providing a robust foundation for training. Experimental results demonstrate that Structural Video Diffusion achieves superior performance in generating lifelike, coherent videos for multiple subjects with dynamic and rich interactions, advancing the state of human-centric video generation. Code is available at https://github.com/zhenzhiwang/Multi-HumanVid
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。