用动态潜变量让虚拟人偶的衣袖自然摆动,无需物理模拟或服装模板。
Latent Dynamics for Full Body Avatar Animation

- 引入可演化潜变量,结合历史姿态预测衣物动态变化
- 在9个真实场景中实现更清晰、连贯的动画效果,视觉质量显著提升
- 适合做高质量虚拟角色动画的研究者和开发者
基于姿态驱动的全身体渲染虚拟人偶虽能生成高质量新视角图像,但松散衣物等动态元素的形变难以仅靠姿态解释:相同姿态可能对应多种状态,其运动依赖历史、惯性与接触。现有方法需专用服装模板或运行时物理模拟,成本高。另一类数据驱动方法虽避免显式分层,但仅通过辅助潜变量建模变化,未显式刻画潜变量随时间演化的机制。此外,现有架构常难以捕捉细粒度细节,导致渲染模糊与时间伪影。本文在姿态条件3D高斯虚拟人基础上,引入基于Transformer的解码器和一个动力学残差潜变量,该潜变量捕捉超出驱动信号的时序外观与几何变化。推理时,学习到的动力学模型基于短时姿态历史与前一潜变量状态,演化当前潜变量。模型将每次更新分解为驱动、恢复与耗散力,实现低开销、依赖历史的连贯轨迹生成。不同初始条件产生多样且合理的运动路径,力分解也提供如刚度等可调控参数。在九段日常动作与多样松垮服饰的真实采集序列上,定量指标与感知用户实验均显示优于近期数据驱动基线。
原文摘要 · Abstract (English)
Pose-driven full-body avatars built on neural rendering produce high-quality novel views of a captured subject. Yet loose clothing and other dynamic elements deform in ways pose alone cannot explain: the same pose can correspond to many different states, because their motion depends on history, inertia, and contact. Explicit simulation and layered-garment methods can model such dynamics, but they require either a dedicated garment template, which raw multi-view capture does not naturally provide, or a test-time physics simulator with non-trivial runtime cost. A parallel line of work learns data-driven clothing avatars that avoid explicit garment layers. These methods add an auxiliary latent for variation beyond pose; at inference, they fix it, regress it from pose, or retrieve it from training data, without explicitly modeling how the latent evolves with its own dynamics. Additionally, even in everyday motion with loose clothing, existing architectures often struggle to capture fine-grained detail, producing blurry renderings and temporal artifacts. We augment a pose-conditioned 3D Gaussian avatar with a transformer-based decoder and a dynamics residual latent that captures temporal appearance and geometry variation beyond the driving signals. At inference, a learned latent dynamics model evolves the residual latent from a short pose history and the previous latent state. The model decomposes each update into driving, restoring, and dissipative forces, producing temporally coherent, history-dependent rollouts with negligible added cost. Different initial conditions yield diverse yet plausible motion trajectories, and the force decomposition exposes controls such as stiffness. Across nine captured sequences of everyday motion with diverse loose garments, quantitative metrics and a perceptual user study show improved animation quality over recent data-driven baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。