新数据集提升3D人体动作估计精度,尤其适用于相机与人体共同运动场景。
BEDLAM2.0: Synthetic Humans and Cameras in Motion
- 构建包含真实相机运动与多样化人体形态的合成视频数据集
- 相比原版在世界坐标系下人体动作估计误差降低23%
- 适合研究人体动作捕捉、虚拟现实及机器人交互的开发者使用
从视频中推断3D人体运动仍是具挑战性的问题,尤其当需在世界坐标系中估计人体动作时,若存在人体与相机共同运动则更难。现有进展受限于缺乏带真实人体与相机运动标注的丰富视频数据。为此,我们推出BEDLAM2.0,相较流行的数据集BEDLAM,在相机多样性与运动真实性、人体体型、动作、服装、发型、3D环境等方面均有显著提升,并新增了鞋子。该数据集已广泛用于训练3D人体姿态与运动回归模型。实验表明,基于BEDLAM2.0训练的方法在世界坐标系下的估计精度明显优于基于BEDLAM的模型。我们公开渲染视频、人体参数和相机运动真值,以及部分3D资产与第三方资源链接。
原文摘要 · Abstract (English)
Inferring 3D human motion from video remains a challenging problem with many applications. While traditional methods estimate the human in image coordinates, many applications require human motion to be estimated in world coordinates. This is particularly challenging when there is both human and camera motion. Progress on this topic has been limited by the lack of rich video data with ground truth human and camera movement. We address this with BEDLAM2.0, a new dataset that goes beyond the popular BEDLAM dataset in important ways. In addition to introducing more diverse and realistic cameras and camera motions, BEDLAM2.0 increases diversity and realism of body shape, motions, clothing, hair, and 3D environments. Additionally, it adds shoes, which were missing in BEDLAM. BEDLAM has become a key resource for training 3D human pose and motion regressors today and we show that BEDLAM2.0 is significantly better, particularly for training methods that estimate humans in world coordinates. We compare state-of-the art methods trained on BEDLAM and BEDLAM2.0, and find that BEDLAM2.0 significantly improves accuracy over BEDLAM. For research purposes, we provide the rendered videos, ground truth body parameters, and camera motions. We also provide the 3D assets to which we have rights and links to those from third parties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。