动态校准+感知,让自动驾驶卡车在弯道中依然看得准。
Mind the Hitch: Dynamic Calibration and Articulated Perception for Autonomous Trucks
- 用时序注意力融合多视角图像,实时估计挂车与车头的6自由度相对位姿
- 在快速转向和遮挡下仍保持感知准确率,3D检测性能提升12.7%
- 适合做重卡自动驾驶的算法研究者,尤其关注动态校准与复杂结构感知
自动驾驶卡车面临挂车-车头铰接结构带来的独特挑战,以及因第五轮连接点和挂车柔性导致的传感器位姿动态变化。现有感知与标定方法依赖静态基准或高视差纹理丰富的场景,在真实环境下可靠性不足。本文提出dCAP(动态校准与铰接感知)框架,基于视觉连续估计车头与挂车摄像头间的6-DoF相对位姿。dCAP采用结合跨视图与时间注意力的Transformer,有效聚合空间线索并保持时序一致性,实现快速铰接与遮挡下的精准感知。集成于BEVFormer后,以动态预测外参替代静态标定,显著提升3D目标检测性能。为评估引入STT4AT——一个基于CARLA的基准数据集,模拟半挂车在多种环境下的多传感器同步与随时间变化的刚性结构。实验表明,dCAP在复杂工况下实现稳定高精度感知,克服了静态标定的局限。相关数据集、开发工具包及源码将公开发布。
原文摘要 · Abstract (English)
Autonomous trucking poses unique challenges due to articulated tractor-trailer geometry, and time-varying sensor poses caused by the fifth-wheel joint and trailer flex. Existing perception and calibration methods assume static baselines or rely on high-parallax and texture-rich scenes, limiting their reliability under real-world settings. We propose dCAP (dynamic Calibration and Articulated Perception), a vision-based framework that continuously estimates the 6-DoF (degree of freedom) relative pose between tractor and trailer cameras. dCAP employs a transformer with cross-view and temporal attention to robustly aggregate spatial cues while maintaining temporal consistency, enabling accurate perception under rapid articulation and occlusion. Integrated with BEVFormer, dCAP improves 3D object detection by replacing static calibration with dynamically predicted extrinsics. To facilitate evaluation, we introduce STT4AT, a CARLA-based benchmark simulating semi-trailer trucks with synchronized multi-sensor suites and time-varying inter-rig geometry across diverse environments. Experiments demonstrate that dCAP achieves stable, accurate perception while addressing the limitations of static calibration in autonomous trucking. The dataset, development kit, and source code will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。