实时从第一人称视角生成高保真人体动作,支持身份控制。
REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioning
- 采用级联去噪扩散模型,快速建模身体与手部动作关联。
- 单步去噪实现实时推理,精度显著优于现有方法。
- 基于少量动作样本的身份条件化,提升个性化动作生成质量。
我们提出 REWIND(实时第一人称全身动作扩散模型),一种基于第一人称图像输入的实时、高保真人体动作估计方法。现有方法因基于扩散模型的迭代优化导致非实时且非因果,REWIND 则实现完全因果与实时推理。为此,我们引入:(1) 级联身体-手部去噪扩散,以快速前馈方式建模第一人称视角下身体与手部动作的相关性;(2) 扩散蒸馏,使单次去噪即可实现高质量动作估计。模型基于改进的 Transformer 架构,可因果建模输出动作,并增强对未见动作长度的泛化能力。此外,当有身份先验时,可选支持身份条件化动作估计。我们提出一种基于目标身份少量姿态样例的新颖条件化方法,进一步提升估计质量。大量实验表明,无论是有无样例条件,REWIND 均显著优于现有基线。
原文摘要 · Abstract (English)
We present REWIND (Real-Time Egocentric Whole-Body Motion Diffusion), a one-step diffusion model for real-time, high-fidelity human motion estimation from egocentric image inputs. While an existing method for egocentric whole-body (i.e., body and hands) motion estimation is non-real-time and acausal due to diffusion-based iterative motion refinement to capture correlations between body and hand poses, REWIND operates in a fully causal and real-time manner. To enable real-time inference, we introduce (1) cascaded body-hand denoising diffusion, which effectively models the correlation between egocentric body and hand motions in a fast, feed-forward manner, and (2) diffusion distillation, which enables high-quality motion estimation with a single denoising step. Our denoising diffusion model is based on a modified Transformer architecture, designed to causally model output motions while enhancing generalizability to unseen motion lengths. Additionally, REWIND optionally supports identity-conditioned motion estimation when identity prior is available. To this end, we propose a novel identity conditioning method based on a small set of pose exemplars of the target identity, which further enhances motion estimation quality. Through extensive experiments, we demonstrate that REWIND significantly outperforms the existing baselines both with and without exemplar-based identity conditioning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。