arXiv:2510.20158cs.CV2025-10被引 2

单目图像下精准估计自行车与骑手的8维姿态,提升自动驾驶安全

Monocular Visual 8D Pose Estimation for Articulated Bicycles and Cyclists

  • 基于单张图像联合预测8维姿态与3D关键点
  • 首次实现对车把和踏板角度的精确估计
  • 适合自动驾驶中行人与骑行者行为预测场景

在自动驾驶中,骑行者属于高危道路使用者(VRU),其姿态准确估计对横穿意图识别、行为预测与避撞至关重要。传统6D姿态方法仅适用于刚性物体,而可动自行车由多个通过关节连接的刚性部件组成,其3D边界框随车把、踏板角度变化而改变,且方向不必然与车把朝向一致。本文提出一种类别级8D姿态估计方法,从单张RGB图像中同时估计自行车的3D平移、旋转,以及车把和踏板相对于车身的旋转角。该方法联合估计8D姿态与3D关键点,并在合成与真实图像混合数据上训练,具有良好泛化能力。实验表明,本方法在8D姿态估计精度上优于使用刚性模板的最新6D姿态估计算法。

原文摘要 · Abstract (English)

In Autonomous Driving, cyclists belong to the safety-critical class of Vulnerable Road Users (VRU), and accurate estimation of their pose is critical for cyclist crossing intention classification, behavior prediction, and collision avoidance. Unlike rigid objects, articulated bicycles are composed of movable rigid parts linked by joints and constrained by a kinematic structure. 6D pose methods can estimate the 3D rotation and translation of rigid bicycles, but 6D becomes insufficient when the steering/pedals angles of the bicycle vary. That is because: 1) varying the articulated pose of the bicycle causes its 3D bounding box to vary as well, and 2) the 3D box orientation is not necessarily aligned to the orientation of the steering which determines the actual intended travel direction. In this work, we introduce a method for category-level 8D pose estimation for articulated bicycles and cyclists from a single RGB image. Besides being able to estimate the 3D translation and rotation of a bicycle from a single image, our method also estimates the rotations of its steering handles and pedals with respect to the bicycle body frame. These two new parameters enable the estimation of a more fine-grained bicycle pose state and travel direction. Our proposed model jointly estimates the 8D pose and the 3D Keypoints of articulated bicycles, and trains with a mix of synthetic and real image data to generalize on real images. We include an evaluation section where we evaluate the accuracy of our estimated 8D pose parameters, and our method shows promising results by achieving competitive scores when compared against state-of-the-art category-level 6D pose estimators that use rigid canonical object templates for matching.

姿态估计自动驾驶单目视觉自行车

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。