可控自由视角3D人物视频生成,支持自定义身份、动作与场景。
CFSynthesis: Controllable and Free-view 3D Human Video Synthesis
- 基于纹理-SMPL表示,保持多视角下角色外观一致。
- 分离前景与背景,实现用户指定场景的无缝融合。
- 在复杂动作和自由视角上表现优异,适合影视创作应用。
人体视频合成旨在生成逼真的人物形象,广泛应用于VR、故事讲述和内容创作。尽管基于2D扩散的方法取得显著进展,但在复杂3D姿态和多变场景背景下仍存在泛化困难。为此,我们提出CFSynthesis,一种可生成高质量人体视频的新框架,支持自定义身份、运动及场景配置。该方法采用纹理-SMPL表示,确保自由视角下的角色外观一致性;同时引入新颖的前景-背景分离策略,有效分解场景并实现用户指定背景的无缝融合。在多个数据集上的实验表明,CFSynthesis不仅在复杂人体动画中达到领先性能,还能在自由视角和用户指定场景中良好适应。
原文摘要 · Abstract (English)
Human video synthesis aims to create lifelike characters in various environments, with wide applications in VR, storytelling, and content creation. While 2D diffusion-based methods have made significant progress, they struggle to generalize to complex 3D poses and varying scene backgrounds. To address these limitations, we introduce CFSynthesis, a novel framework for generating high-quality human videos with customizable attributes, including identity, motion, and scene configurations. Our method leverages a texture-SMPL-based representation to ensure consistent and stable character appearances across free viewpoints. Additionally, we introduce a novel foreground-background separation strategy that effectively decomposes the scene as foreground and background, enabling seamless integration of user-defined backgrounds. Experimental results on multiple datasets show that CFSynthesis not only achieves state-of-the-art performance in complex human animations but also adapts effectively to 3D motions in free-view and user-specified scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。