融合事件相机与视觉信息,实现高速移动物体的精准三维姿态估计
PoseStreamer: A Multi-modal Framework for 3D Tracking of Unseen Moving Objects
- 采用多模态融合架构,结合历史姿态与2D跟踪先验提升稳定性
- 在高速运动场景下3D中心定位召回率显著提升,精度优于现有方法
- 无需模板,适用于未知物体,适合机器人视觉与自动驾驶应用
六自由度(6DoF)姿态估计在计算机视觉中至关重要,但在高速和低光条件下,传统RGB相机易受运动模糊影响。事件相机虽具高时间分辨率,但现有6DoF方法在高速运动时性能仍不理想。为此,我们提出PoseStreamer,一种专为高速运动设计的鲁棒多模态6DoF姿态估计框架。该框架包含三个核心组件:自适应姿态记忆队列,利用历史方向信息保证时序一致性;以物体为中心的2D追踪器,提供强2D先验以提升3D中心召回率;沿相机射线进行几何精修的射线姿态滤波器。此外,我们构建了新的多模态数据集MoCapCube6D,用于评估快速运动下的性能。大量实验表明,PoseStreamer不仅在高速场景中实现更高精度,且作为无需模板的通用框架,对未见过的运动物体具有强泛化能力。
原文摘要 · Abstract (English)
Six degree of freedom (6DoF) pose estimation for novel objects is a critical task in computer vision, yet it faces significant challenges in high-speed and low-light scenarios where standard RGB cameras suffer from motion blur. While event cameras offer a promising solution due to their high temporal resolution, current 6DoF pose estimation methods typically yield suboptimal performance in high-speed object moving scenarios. To address this gap, we propose PoseStreamer, a robust multi-modal 6DoF pose estimation framework designed specifically on high-speed moving scenarios. Our approach integrates three core components: an Adaptive Pose Memory Queue that utilizes historical orientation cues for temporal consistency, an Object-centric 2D Tracker that provides strong 2D priors to boost 3D center recall, and a Ray Pose Filter for geometric refinement along camera rays. Furthermore, we introduce MoCapCube6D, a novel multi-modal dataset constructed to benchmark performance under rapid motion. Extensive experiments demonstrate that PoseStreamer not only achieves superior accuracy in high-speed moving scenarios, but also exhibits strong generalizability as a template-free framework for unseen moving objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。