用自回归潜变量让3D高斯角色动画更真实稳定
Autoregressive Appearance Prediction for 3D Gaussian Avatars
- 用姿态+外观潜变量控制3D高斯点云,解耦动作与细节
- 自回归预测潜变量,使表情和衣物动态变化平滑连续
- 适合追求高质量、稳定人物动画的开发者或研究者
逼真的沉浸式人物角色体验需要捕捉衣物、头发动态、细微面部表情及个人特有动作模式等精细细节。这通常依赖大规模高质量数据集,但相似姿态可能对应不同外观,导致模型训练时产生歧义和虚假关联。现有模型在训练中拟合这些细节易过拟合,对新姿态生成不稳定、突变的外观。本文提出一种基于空间MLP主干的3D高斯泼溅角色模型,同时依赖姿态和外观潜变量进行条件控制。该潜变量通过编码器在训练中学习,形成紧凑表征,提升重建质量并帮助消除姿态驱动渲染的歧义。推理时,我们的预测器自回归推断潜变量,实现时间上平滑的外观演化,显著改善稳定性。整体方法为高保真、稳定的角色驱动提供了可靠路径。
原文摘要 · Abstract (English)
A photorealistic and immersive human avatar experience demands capturing fine, person-specific details such as cloth and hair dynamics, subtle facial expressions, and characteristic motion patterns. Achieving this requires large, high-quality datasets, which often introduce ambiguities and spurious correlations when very similar poses correspond to different appearances. Models that fit these details during training can overfit and produce unstable, abrupt appearance changes for novel poses. We propose a 3D Gaussian Splatting avatar model with a spatial MLP backbone that is conditioned on both pose and an appearance latent. The latent is learned during training by an encoder, yielding a compact representation that improves reconstruction quality and helps disambiguate pose-driven renderings. At driving time, our predictor autoregressively infers the latent, producing temporally smooth appearance evolution and improved stability. Overall, our method delivers a robust and practical path to high-fidelity, stable avatar driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。