用单视角生成带运动结构的视频,让物体部件自动对齐关节动作。
Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
- 双分支扩散模型同步生成图像和可动部件图
- 20,000个带骨骼绑定的物体数据集支持训练与评估
- 生成部件可轻松转为3D骨架,适合动画制作
我们提出Stable Part Diffusion 4D(SP4D),一种从单视角输入生成配对的RGB与运动部件视频的框架。不同于依赖外观语义的传统部件分割方法,SP4D学习生成与物体运动关节对齐的运动部件,且在不同视角和时间上保持一致。该框架采用双分支扩散模型,联合合成RGB帧与对应的部件分割图。为简化结构并灵活支持不同部件数量,引入空间颜色编码方案,将部件掩码映射为连续的类似RGB图像,使分割分支可共享RGB分支的潜在VAE,同时通过后处理恢复部件分割。双向扩散融合(BiDiFuse)模块增强跨分支一致性,并辅以对比性部件一致性损失,促进部件预测在空间和时间上的对齐。实验表明,生成的2D部件图可被提升至3D,仅需少量人工调整即可获得骨骼结构与谐波蒙皮权重。为训练与评估SP4D,我们构建了KinematicParts20K数据集,从Objaverse XL(Deitke et al., 2023)中筛选并处理超过20,000个带骨骼绑定的物体,每个物体配有多个视角的RGB与部件视频序列。实验显示,SP4D在多样化场景下表现强泛化能力,包括真实视频、新生成物体及罕见关节姿态,生成结果具备运动感知特性,适用于下游动画与运动相关任务。
原文摘要 · Abstract (English)
We present Stable Part Diffusion 4D (SP4D), a framework for generating paired RGB and kinematic part videos from monocular inputs. Unlike conventional part segmentation methods that rely on appearance-based semantic cues, SP4D learns to produce kinematic parts - structural components aligned with object articulation and consistent across views and time. SP4D adopts a dual-branch diffusion model that jointly synthesizes RGB frames and corresponding part segmentation maps. To simplify the architecture and flexibly enable different part counts, we introduce a spatial color encoding scheme that maps part masks to continuous RGB-like images. This encoding allows the segmentation branch to share the latent VAE from the RGB branch, while enabling part segmentation to be recovered via straightforward post-processing. A Bidirectional Diffusion Fusion (BiDiFuse) module enhances cross-branch consistency, supported by a contrastive part consistency loss to promote spatial and temporal alignment of part predictions. We demonstrate that the generated 2D part maps can be lifted to 3D to derive skeletal structures and harmonic skinning weights with few manual adjustments. To train and evaluate SP4D, we construct KinematicParts20K, a curated dataset of over 20K rigged objects selected and processed from Objaverse XL (Deitke et al., 2023), each paired with multi-view RGB and part video sequences. Experiments show that SP4D generalizes strongly to diverse scenarios, including real-world videos, novel generated objects, and rare articulated poses, producing kinematic-aware outputs suitable for downstream animation and motion-related tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。