用骨骼驱动生成可编辑的动态3D高斯,让动作修改更直观。
SkeletonGaussian: Editable 4D Generation through Gaussian Skeletonization
- 通过分层骨骼结构分解刚性与非刚性运动,实现显式控制。
- 在单目视频输入下生成高质量动态3D高斯,支持直接编辑动作。
- 适合需要精细动作调整的动画生成、虚拟人建模场景。
4D生成在从文本、图像或视频合成动态3D物体方面取得了显著进展。然而,现有方法通常将运动表示为隐式变形场,限制了直接控制与可编辑性。为此,我们提出SkeletonGaussian,一种从单目视频输入生成可编辑动态3D高斯的新框架。该方法引入分层关节化表示,将运动分解为由骨骼显式驱动的稀疏刚性运动和细粒度非刚性运动。具体而言,我们提取鲁棒骨骼,并通过线性混合皮肤(linear blend skinning)驱动刚性运动,随后采用基于六面体平面(hexplane)的细化模块处理非刚性形变,显著提升可解释性与可编辑性。实验表明,SkeletonGaussian在生成质量上优于现有方法,同时支持直观的动作编辑,确立了可编辑4D生成的新范式。
原文摘要 · Abstract (English)
4D generation has made remarkable progress in synthesizing dynamic 3D objects from input text, images, or videos. However, existing methods often represent motion as an implicit deformation field, which limits direct control and editability. To address this issue, we propose SkeletonGaussian, a novel framework for generating editable dynamic 3D Gaussians from monocular video input. Our approach introduces a hierarchical articulated representation that decomposes motion into sparse rigid motion explicitly driven by a skeleton and fine-grained non-rigid motion. Concretely, we extract a robust skeleton and drive rigid motion via linear blend skinning, followed by a hexplane-based refinement for non-rigid deformations, enhancing interpretability and editability. Experimental results demonstrate that SkeletonGaussian surpasses existing methods in generation quality while enabling intuitive motion editing, establishing a new paradigm for editable 4D generation. Project page: https://wusar.github.io/projects/skeletongaussian/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。