提出LEGO-Motion框架,融合占用网格与实例特征,提升自动驾驶运动预测的准确性和物理一致性。
LEGO-Motion: Learning-Enhanced Grids with Occupancy Instance Modeling for Class-Agnostic Motion Prediction
- 在鸟瞰图空间中融合占用网格与实例特征,增强环境理解
- 在nuScenes数据集上达到当前最优性能,显著优于现有方法
- 适用于开放场景,适合追求高精度运动预测的研究与工程应用
精确可靠的时空信息对自动驾驶系统至关重要。现有基于物体的感知模型难以应对开放场景类别,且缺乏精确的内在几何表征;而基于占用率的无类别方法虽能全面刻画场景,却难以保证物理一致性,且忽略交通参与者间的交互关系,制约了运动预测的准确性。本文提出一种新的无类别运动预测框架LEGO-Motion,将实例特征融入鸟瞰图(BEV)空间。模型包含三个模块:BEV编码器、交互增强型实例编码器和实例增强型BEV编码器,有效提升了交互建模能力与物理一致性,从而实现更精准稳健的环境理解。在nuScenes数据集上的大量实验表明,本方法性能达当前最优;在先进的FMCW LiDAR基准测试中也验证了其实际应用潜力与泛化能力。代码将公开以促进后续研究。
原文摘要 · Abstract (English)
Accurate and reliable spatial and motion information plays a pivotal role in autonomous driving systems. However, object-level perception models struggle with handling open scenario categories and lack precise intrinsic geometry. On the other hand, occupancy-based class-agnostic methods excel in representing scenes but fail to ensure physics consistency and ignore the importance of interactions between traffic participants, hindering the model's ability to learn accurate and reliable motion. In this paper, we introduce a novel occupancy-instance modeling framework for class-agnostic motion prediction tasks, named LEGO-Motion, which incorporates instance features into Bird's Eye View (BEV) space. Our model comprises (1) a BEV encoder, (2) an Interaction-Augmented Instance Encoder, and (3) an Instance-Enhanced BEV Encoder, improving both interaction relationships and physics consistency within the model, thereby ensuring a more accurate and robust understanding of the environment. Extensive experiments on the nuScenes dataset demonstrate that our method achieves state-of-the-art performance, outperforming existing approaches. Furthermore, the effectiveness of our framework is validated on the advanced FMCW LiDAR benchmark, showcasing its practical applicability and generalization capabilities. The code will be made publicly available to facilitate further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。