arXiv:2508.09404cs.CVcs.MM2025-08被引 4

构建首个高精度3D行人交互动作数据集,助力自动驾驶理解复杂城市场景中人群行为。

Waymo-3DSkelMo: A Multi-Agent 3D Skeletal Motion Dataset for Pedestrian Interaction Modeling in Autonomous Driving

  • 基于激光点云与人体运动先验,提升3D姿态序列质量与时间连续性。
  • 覆盖超14000秒、800+真实驾驶场景,单场景最多250名参与者。
  • 适用于高密度行人环境下的行为预测研究,支持未来智能驾驶系统开发。

大规模高质量的多行人交互3D运动数据集对于自动驾驶中实现精细的行人交互理解至关重要。然而,现有数据集多依赖单目视频估计3D姿态,易受遮挡影响且缺乏时间连续性,导致动作不真实、质量低。本文提出Waymo-3DSkelMo,首个基于Waymo Perception数据集构建的大规模高精度3D骨骼运动数据集,提供具显式交互语义的时序连贯3D姿态序列。核心思想是利用3D人体形状与运动先验,从原始LiDAR点云中优化提取3D姿态。数据集涵盖超过14,000秒、800多个真实驾驶场景,平均每场景27名参与者(最多达250人)。我们建立了不同行人密度下的3D姿态预测基准,结果表明其作为复杂城市环境中精细化人类行为理解的基础资源具有重要价值。数据与代码将公开于https://github.com/GuangxunZhu/Waymo-3DSkelMo。

原文摘要 · Abstract (English)

Large-scale high-quality 3D motion datasets with multi-person interactions are crucial for data-driven models in autonomous driving to achieve fine-grained pedestrian interaction understanding in dynamic urban environments. However, existing datasets mostly rely on estimating 3D poses from monocular RGB video frames, which suffer from occlusion and lack of temporal continuity, thus resulting in unrealistic and low-quality human motion. In this paper, we introduce Waymo-3DSkelMo, the first large-scale dataset providing high-quality, temporally coherent 3D skeletal motions with explicit interaction semantics, derived from the Waymo Perception dataset. Our key insight is to utilize 3D human body shape and motion priors to enhance the quality of the 3D pose sequences extracted from the raw LiDRA point clouds. The dataset covers over 14,000 seconds across more than 800 real driving scenarios, including rich interactions among an average of 27 agents per scene (with up to 250 agents in the largest scene). Furthermore, we establish 3D pose forecasting benchmarks under varying pedestrian densities, and the results demonstrate its value as a foundational resource for future research on fine-grained human behavior understanding in complex urban environments. The dataset and code will be available at https://github.com/GuangxunZhu/Waymo-3DSkelMo

3D姿态行人交互自动驾驶数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。