用视频模型先验构建3D角色的合理动作空间,让自动绑定模型更自然好用。
ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes
- 从预训练视频扩散模型中提取动作先验,生成符合物理规律的关节配置分布。
- 无需人工标注数据,在零样本下对新物种和骨骼结构仍能保持动作合理性。
- 支持动画采样、逆运动学投影与连贯轨迹生成,打通2D生成与3D控制闭环。
运动学绑定为3D网格提供结构化操控接口,但缺乏对应的姿态空间——即针对特定网格的合理关节配置显式表示。缺少姿态空间会导致随机采样或手动调整参数时产生语义或几何错误,如解剖超伸展和非物理自交。我们提出视频引导的姿态空间(ViPS),一种前馈框架,通过蒸馏预训练视频扩散模型的动作先验,发现自动绑定网格的有效动作潜在分布。不同于依赖稀缺艺术家创作4D数据集或仅重建单个动作实例的方法,ViPS将生成式视频模型先验转化为通用的姿态分布。应用于蒙皮网格的可微分几何验证器确保形状特异性完整性,无需人工正则项。我们的前馈模型揭示了平滑、紧凑且可控的姿态空间,支持多样化形状采样、流形投影实现逆运动学求解,以及动画与关键帧所需的时序连贯轨迹。此外,蒸馏出的3D姿态样本作为语义代理,可引导视频扩散模型,有效实现生成2D先验与结构化3D运动控制之间的闭环。评估表明,仅使用视频先验训练的ViPS,在合理性和多样性上达到与基于合成艺术家创建4D数据训练的顶尖模型相当的性能。同时,作为通用模型,ViPS在未见物种和骨骼拓扑上表现出稳健的零样本泛化能力。
原文摘要 · Abstract (English)
Kinematic rigs provide a structured interface for articulating 3D meshes but lack any associated pose space, i.e., an explicit representation of the plausible manifold of joint configurations for a given mesh. Without such a pose space, stochastic sampling or manual manipulation of raw rig parameters easily results in semantic and/or geometric violations, such as anatomical hyperextension and non-physical self-intersections. We propose Video-informed Pose Spaces (ViPS), a feedforward framework that discovers the latent distribution of valid articulations for auto-rigged meshes by distilling motion priors from a pretrained video diffusion model. Unlike existing methods that rely on scarce, artist-authored 4D datasets, or focus on reconstructing instances of individual motions, ViPS transfers generative video model priors into a universal distribution over the given rig parameterization. Differentiable geometric validators applied to the skinned mesh enforce shape-specific integrity without requiring manual regularizers. Our feedforward model reveals a smooth, compact, and controllable pose space. This, in turn, supports sampling for diverse shape variations, manifold projection for inverse kinematics, and temporally coherent trajectories for animation and keyframing. Further, the distilled 3D pose samples serve as semantic proxies to guide video diffusion, effectively closing the loop between generative 2D priors and structured 3D kinematic control. Our evaluations show that ViPS, trained solely using video priors, matches the performance of state-of-the-art models trained on synthetic artist-created 4D data in both plausibility and diversity. Additionally, as a universal model, ViPS exhibits robust zero-shot generalization to out-of-distribution species and unseen skeletal topologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。