用稀疏单目视频生成高保真3D avatar动画,提升表情控制精度
LiftAvatar: Kinematic-Space Completion for Expression-Controlled 3D Gaussian Avatar Animation

- 在运动空间补全稀疏输入,生成丰富姿态-表情组合
- 多参考条件聚合实现3D一致性与强可控性,显著减少重建伪影
- 可即插即用,适合需要高精度表情驱动的3D avatar应用
我们提出LiftAvatar,一种新范式,在运动空间(如面部表情和头部姿态)中补全稀疏的单目观测,并利用补全信号驱动高保真3D avatar动画。LiftAvatar是一种细粒度、表情可控的大规模视频扩散Transformer,能够基于单张或多张参考图像合成高质量、时间连贯的表情序列。核心思想是将不完整的输入数据提升为更丰富的运动表示,从而增强下游3D avatar流水线中的重建与动画效果。为此,我们引入:(i) 多粒度表情控制机制,结合阴影图与表情系数实现精确稳定驱动;(ii) 多参考条件机制,聚合多帧互补线索,实现强3D一致性和可控性。作为即插即用增强器,LiftAvatar直接解决基于3D高斯溅射的avatar因日常单目视频中运动线索稀疏导致的表现力不足与重建伪影问题。通过扩展不完整观测为多样姿态-表情变化,还实现了从大规模视频生成模型向3D流水线的有效先验蒸馏,带来显著性能提升。大量实验表明,LiftAvatar持续提升当前先进3D avatar方法的动画质量与定量指标,尤其在极端、未见过的表情下表现突出。
原文摘要 · Abstract (English)
We present LiftAvatar, a new paradigm that completes sparse monocular observations in kinematic space (e.g., facial expressions and head pose) and uses the completed signals to drive high-fidelity avatar animation. LiftAvatar is a fine-grained, expression-controllable large-scale video diffusion Transformer that synthesizes high-quality, temporally coherent expression sequences conditioned on single or multiple reference images. The key idea is to lift incomplete input data into a richer kinematic representation, thereby strengthening both reconstruction and animation in downstream 3D avatar pipelines. To this end, we introduce (i) a multi-granularity expression control scheme that combines shading maps with expression coefficients for precise and stable driving, and (ii) a multi-reference conditioning mechanism that aggregates complementary cues from multiple frames, enabling strong 3D consistency and controllability. As a plug-and-play enhancer, LiftAvatar directly addresses the limited expressiveness and reconstruction artifacts of 3D Gaussian Splatting-based avatars caused by sparse kinematic cues in everyday monocular videos. By expanding incomplete observations into diverse pose-expression variations, LiftAvatar also enables effective prior distillation from large-scale video generative models into 3D pipelines, leading to substantial gains. Extensive experiments show that LiftAvatar consistently boosts animation quality and quantitative metrics of state-of-the-art 3D avatar methods, especially under extreme, unseen expressions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。