用图注意力网络融合多传感器数据,实现自动驾驶中动态目标的高精度形状与轨迹追踪。
LEO: Graph Attention Network based Hybrid Multi Sensor Extended Object Fusion and Tracking for Autonomous Driving Applications
- 基于图注意力网络融合多模态传感器轨迹,自适应学习融合权重。
- 在奔驰Drive Pilot数据集上实现实时计算,长距离目标追踪误差低于0.5米。
- 适用于复杂车辆如挂车,跨数据集泛化能力强,适合量产系统部署。
动态物体的精确形状与轨迹估计对可靠自动驾驶至关重要。传统贝叶斯扩展目标模型虽理论稳健高效,但依赖先验与似然函数的完备性;深度学习方法具适应性,却需密集标注且计算开销大。本文提出LEO(Learned Extension of Objects),一种时空图注意力网络,融合多模态量产级传感器轨迹,自动学习融合权重,保证时序一致性,并表征多尺度形状。采用任务特异的平行四边形真值格式,能够建模复杂几何形态(如铰接式卡车和挂车),并跨传感器类型、配置、目标类别与区域实现泛化,对远距离及挑战性目标仍保持鲁棒性。在梅赛德斯-奔驰 DRIVE PILOT SAE L3 数据集上的评估表明其具备生产系统所需的实时计算效率;在公开数据集 View of Delft (VoD) 上的额外验证进一步证实其跨数据集泛化能力。
原文摘要 · Abstract (English)
Accurate shape and trajectory estimation of dynamic objects is essential for reliable automated driving. Classical Bayesian extended-object models offer theoretical robustness and efficiency but depend on completeness of a-priori and update-likelihood functions, while deep learning methods bring adaptability at the cost of dense annotations and high compute. We bridge these strengths with LEO (Learned Extension of Objects), a spatio-temporal Graph Attention Network that fuses multi-modal production-grade sensor tracks to learn adaptive fusion weights, ensure temporal consistency, and represent multi-scale shapes. Using a task-specific parallelogram ground-truth formulation, LEO models complex geometries (e.g. articulated trucks and trailers) and generalizes across sensor types, configurations, object classes, and regions, remaining robust for challenging and long-range targets. Evaluations on the Mercedes-Benz DRIVE PILOT SAE L3 dataset demonstrate real-time computational efficiency suitable for production systems; additional validation on public datasets such as View of Delft (VoD) further confirms cross-dataset generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。