为机器人应用提供复杂真实场景下的多人3D姿态数据集。
JRDB-Pose3D: A Multi-person 3D Human Pose and Shape Estimation Dataset for Robotics
- 从移动机器人平台采集多人群动态场景,支持长时间跟踪。
- 每帧平均5-10人,最多单帧35人,含遮挡与肢体截断等真实挑战。
- 适合研究机器人感知、人机交互及社会行为理解的团队使用。
现实场景通常拥挤复杂。因此,准确估计所有附近人员的3D姿态,长期追踪其运动,并理解其在社交与环境背景中的行为,对自动驾驶、机器人感知、导航及人机交互等应用至关重要。然而,现有3D人体姿态数据集大多聚焦单人场景或在受控实验室中采集,难以反映真实环境。为此,我们提出JRDB-Pose3D,该数据集通过移动机器人平台捕获室内与室外多人群体环境,提供丰富的3D人体姿态标注,包括基于SMPL的姿势参数与一致的体形参数,并为每个个体提供时间连续的轨迹ID。JRDB-Pose3D平均每帧包含5-10个3D人体姿态,部分场景同时出现多达35人。数据集面临频繁遮挡、身体截断和出框肢体等真实挑战。此外,它继承了原JRDB数据集全部标注信息,包括2D姿态、社交分组、活动与交互信息、全场景语义掩码、人与物体级别的持续跟踪,以及年龄、性别、种族等个体属性标注,是面向下游感知与以人为中心理解任务的综合性数据集。
原文摘要 · Abstract (English)
Real-world scenes are inherently crowded. Hence, estimating 3D poses of all nearby humans, tracking their movements over time, and understanding their activities within social and environmental contexts are essential for many applications, such as autonomous driving, robot perception, robot navigation, and human-robot interaction. However, most existing 3D human pose estimation datasets primarily focus on single-person scenes or are collected in controlled laboratory environments, which restricts their relevance to real-world applications. To bridge this gap, we introduce JRDB-Pose3D, which captures multi-human indoor and outdoor environments from a mobile robotic platform. JRDB-Pose3D provides rich 3D human pose annotations for such complex and dynamic scenes, including SMPL-based pose annotations with consistent body-shape parameters and track IDs for each individual over time. JRDB-Pose3D contains, on average, 5-10 human poses per frame, with some scenes featuring up to 35 individuals simultaneously. The proposed dataset presents unique challenges, including frequent occlusions, truncated bodies, and out-of-frame body parts, which closely reflect real-world environments. Moreover, JRDB-Pose3D inherits all available annotations from the JRDB dataset, such as 2D pose, information about social grouping, activities, and interactions, full-scene semantic masks with consistent human- and object-level tracking, and detailed annotations for each individual, such as age, gender, and race, making it a holistic dataset for a wide range of downstream perception and human-centric understanding tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。