用3D高斯头像实现可控姿态迁移,生成逼真人体动作视频。
Synthetic Human Action Video Data Generation with Pose Transfer
- 基于3D高斯avatar进行姿态迁移,生成自然人体动作视频。
- 在丰田智能家居和NTU RGB+D数据集上提升动作识别准确率。
- 可扩展小样本数据,补充稀缺人群与多样化背景。
在视频理解任务中,特别是涉及人体运动的任务,合成数据常因诡异特征而效果不佳,导致手语翻译、手势识别及自动驾驶中的人体动作理解难以充分发挥合成数据潜力。本文提出一种基于姿态迁移(具体为可控3D高斯头像模型)的合成人体动作视频生成方法。我们在Toyota Smarthome和NTU RGB+D数据集上进行了评估,结果表明该方法显著提升了动作识别性能。此外,实验还证明该方法能有效扩展少样本数据集,弥补真实训练数据中代表性不足的群体,并增加多样化的背景。本文开源了该方法及RANDOM People数据集,包含从互联网众包获取的新型人类身份视频与头像,用于姿态迁移。
原文摘要 · Abstract (English)
In video understanding tasks, particularly those involving human motion, synthetic data generation often suffers from uncanny features, diminishing its effectiveness for training. Tasks such as sign language translation, gesture recognition, and human motion understanding in autonomous driving have thus been unable to exploit the full potential of synthetic data. This paper proposes a method for generating synthetic human action video data using pose transfer (specifically, controllable 3D Gaussian avatar models). We evaluate this method on the Toyota Smarthome and NTU RGB+D datasets and show that it improves performance in action recognition tasks. Moreover, we demonstrate that the method can effectively scale few-shot datasets, making up for groups underrepresented in the real training data and adding diverse backgrounds. We open-source the method along with RANDOM People, a dataset with videos and avatars of novel human identities for pose transfer crowd-sourced from the internet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。