无需皮肤绑定,直接从2D图像生成3D人物动画。
LUNA: Learning Universal 3D Human Animation Beyond Skinning

- 用Transformer分离全局运动与局部动态,实现精准动作建模。
- 在未标注视频上训练,仍保持高质量3D形变效果。
- 支持零样本跨身份动画,适配图像/草图等多种输入。
从单目图像创建逼真可动画的3D人类形象,仍主要依赖线性混合皮肤(LBS)和参数化人体模型,限制了表现力并常引入拟合误差。我们提出LUNA,一种无LBS的通用神经动画模型,能直接将多模态2D控制信号(如图像、关键点、草图及未见角色)映射为3D高斯形变,跳过显式身体拟合。核心是一个基于Transformer的运动回归器,解耦全局刚性运动与精细局部动态,以捕捉整体运动连贯性与细微非刚性变化。为解决2D到3D提升中的固有歧义并扩展至未拟合数据集,引入混合监督机制:从LBS教师模型中蒸馏软结构先验,并设计损失函数,支持在有限拟合数据和大量野外未标注视频上联合训练。大量实验表明,LUNA在视觉保真度上媲美传统LBS方法,同时实现真实的人体运动与零样本跨身份泛化,适用于多种驱动模态。据我们所知,LUNA是首个支持隐式2D驱动的端到端可动画3D模型。
原文摘要 · Abstract (English)
Creating photorealistic, animatable 3D human avatars from monocular images still largely depends on Linear Blend Skinning (LBS) and parametric body models, which constrain expressivity and often introduce artifacts due to imperfect fitting. We propose LUNA, an LBS-free universal neural animation model that directly maps multiple 2D controls like images, keypoints, sketches, and unseen characters into 3D Gaussian deformations, bypassing explicit body fitting. At its core, a transformer-based motion regressor disentangles global rigid motion from fine-grained local dynamics to capture both coherent movement and subtle non-rigid effects. To resolve the inherent ambiguity of 2D-to-3D lifting while scaling beyond fitted datasets, we introduce hybrid supervision that distills soft structural priors from an LBS teacher and a loss that supports training on both limited fitted data and large in-the-wild unlabeled videos. Extensive experiments show LUNA achieves competitive visual fidelity compared to LBS-based approaches, while delivering realistic human motion and zero-shot cross-identity generalization across diverse driving modalities. To the best of our knowledge, LUNA is the first end-to-end 3D animatable model that supports implicit 2D driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。