融合视觉、惯性与定位数据,实现无硬件同步的高精度姿态估计
DAP-Pose: Deep Temporal Alignment and Physics-aware Cross-modal Sensor Fusion for Robust Pose Estimation

- 分层跨模态融合捕捉视觉与惯性信号的互补信息
- 隐空间对齐解决异步数据时间偏移,误差低于1.31%和0.46度
- 引入物理约束提升稳定性,适合复杂环境下的自动驾驶系统
多模态传感器在复杂环境下实现鲁棒且精确的姿态估计是自动驾驶与移动机器人系统的基础。本文提出DAP-Pose,一种统一的端到端多模态姿态估计模型。该模型引入双层次跨模态融合(BCF)模块,从视觉、惯性及GNSS测量中捕获互补的语义与几何运动线索。为处理时间偏移,设计深度时间对齐(DTA)模块,在隐空间显式对齐异步数据流,实现无需严格硬件同步的连贯运动建模。此外,通过流形几何与GNSS引导的绝对尺度约束引入物理感知,强化运动一致性并抑制漂移。在公开的KITTI基准数据集上进行实验,DAP-Pose达到当前最优性能:平均相对平移误差($t_{rel}$)为1.31%,旋转误差($r_{rel}$)为0.46°。同时,在人为注入严重时间错位条件下仍能准确估计姿态并保持鲁棒性。
原文摘要 · Abstract (English)
Robust and accurate pose estimation with multi-modal sensors is fundamental for autonomous vehicles and mobile robotic systems in complex environments. In this paper, we propose DAP-Pose, a unified end-to-end model for robust multi-modal pose estimation. DAP-Pose introduces a Bi-level Cross-modal Fusion (BCF) module that captures complementary semantic and geometric motion cues from visual, inertial, and GNSS measurements. To handle temporal offsets, we designed a Deep Temporal Alignment (DTA) module that explicitly aligns asynchronous streams in latent space, enabling coherent motion modeling without strict hardware synchronization. Furthermore, we incorporate physics-aware constraints via manifold geometry and GNSS-guided absolute metric scale, enforcing motion consistency and mitigating drift. Experiments upon the public KITTI benchmark dataset were conducted to evaluate the performance of DAP-Pose against existing methods. DAP-Pose achieved the state-of-the-art performance, with the lowest average translation error ($t_{rel}$) of 1.31% and rotation error ($r_{rel}$) of 0.46$^{\circ}$. Furthermore, it accurately estimates poses and maintains robust performance under severe artificially injected temporal misalignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。