用扩散模型补全视频中缺失的人体动作,让虚拟人能自由生成新动作。
Diffusion Model-based Activity Completion for AI Motion Capture from Videos
- 基于扩散模型生成动作间缺失的过渡序列,保持运动连贯性。
- 在Human3.6M数据集上,ADE、FDE和MMADE指标优于现有方法。
- 模型更轻量(16.84M),生成动作更自然,还支持提取传感器数据。
基于AI的动作捕捉技术为传统系统提供了低成本替代方案,但现有方法完全依赖观测视频,要求所有动作预先定义,无法生成未出现的动作。为解决此问题,我们提出一种基于扩散模型的动作补全方法,用于虚拟人生成超出训练数据的动作。假设训练数据包含大量动作片段,但片段间的过渡缺失。通过引入门控模块与时空嵌入模块,我们的方法在Human3.6M数据集上实现优异表现:MDC-Net在ADE、FDE和MMADE指标上优于现有方法,仅在MMFDE上略逊;模型规模仅为16.84M,远小于HumanMAC的28.40M;且生成动作更自然连贯。此外,我们还提出从运动序列中提取加速度与角速度等传感器数据的方法。
原文摘要 · Abstract (English)
AI-based motion capture is an emerging technology that offers a cost-effective alternative to traditional motion capture systems. However, current AI motion capture methods rely entirely on observed video sequences, similar to conventional motion capture. This means that all human actions must be predefined, and movements outside the observed sequences are not possible. To address this limitation, we aim to apply AI motion capture to virtual humans, where flexible actions beyond the observed sequences are required. We assume that while many action fragments exist in the training data, the transitions between them may be missing. To bridge these gaps, we propose a diffusion-model-based action completion technique that generates complementary human motion sequences, ensuring smooth and continuous movements. By introducing a gate module and a position-time embedding module, our approach achieves competitive results on the Human3.6M dataset. Our experimental results show that (1) MDC-Net outperforms existing methods in ADE, FDE, and MMADE but is slightly less accurate in MMFDE, (2) MDC-Net has a smaller model size (16.84M) compared to HumanMAC (28.40M), and (3) MDC-Net generates more natural and coherent motion sequences. Additionally, we propose a method for extracting sensor data, including acceleration and angular velocity, from human motion sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。