用游戏引擎生成百万级带真实标注的3D人体姿态数据,解决真实数据难获取问题
UnrealPose: Leveraging Game Engine Kinematics for Large-Scale Synthetic Human Pose Data
- 基于虚幻引擎5构建合成数据生成管道,自动输出多类标注信息
- 生成100万帧数据,覆盖5个场景、100种动作、5名角色,视角多样
- 可直接用于姿态估计、检测等任务,适合研究数据增强与模型训练
多样且精确标注的3D人体姿态数据成本高且依赖专业工作室,而真实世界数据缺乏已知真值。我们提出UnrealPose-Gen,一个基于虚幻引擎5和Movie Render Queue的离线渲染管线,可生成高质量图像并附带:(i) 世界坐标与相机坐标下的3D关节位置,(ii) 2D投影点、COCO风格关键点及遮挡/可见性标志,(iii) 人物边界框,(iv) 相机内参与外参。利用该管道构建了UnrealPose-1M数据集,包含约一百万帧,涵盖八段序列:五段脚本化“连贯”序列(五个场景,约40种动作,五名角色),三段随机序列(三个场景,约100种动作,五名角色),所有数据均从多样化相机轨迹采集,实现广泛视角覆盖。通过四项任务验证真实性:图像到3D姿态、2D关键点检测、2D到3D提升、人物检测/分割。尽管时间和资源限制无法无限扩展,我们仍公开发布UnrealPose-1M数据集及UnrealPose-Gen生成管道,支持第三方自主生成人体姿态数据。
原文摘要 · Abstract (English)
Diverse, accurately labeled 3D human pose data is expensive and studio-bound, while in-the-wild datasets lack known ground truth. We introduce UnrealPose-Gen, an Unreal Engine 5 pipeline built on Movie Render Queue for high-quality offline rendering. Our generated frames include: (i) 3D joints in world and camera coordinates, (ii) 2D projections and COCO-style keypoints with occlusion and joint-visibility flags, (iii) person bounding boxes, and (iv) camera intrinsics and extrinsics. We use UnrealPose-Gen to present UnrealPose-1M, an approximately one million frame corpus comprising eight sequences: five scripted "coherent" sequences spanning five scenes, approximately 40 actions, and five subjects; and three randomized sequences across three scenes, approximately 100 actions, and five subjects, all captured from diverse camera trajectories for broad viewpoint coverage. As a fidelity check, we report real-to-synthetic results on four tasks: image-to-3D pose, 2D keypoint detection, 2D-to-3D lifting, and person detection/segmentation. Though time and resources constrain us from an unlimited dataset, we release the UnrealPose-1M dataset, as well as the UnrealPose-Gen pipeline to support third-party generation of human pose data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。