arXiv:2508.07089cs.CVcs.RO2025-08ICCV被引 5

端到端联合检测与轨迹预测,提升自动驾驶感知精度

ForeSight: Multi-View Streaming Joint Object Detection and Trajectory Forecasting

  • 采用双向流式学习共享查询记忆,实现检测与预测协同优化
  • 在nuScenes上达54.9% EPA,较前方法提升9.3个百分点
  • 无需显式目标关联,降低误差传播,适合多帧序列实时应用

我们提出ForeSight,一种面向自动驾驶视觉3D感知的新型联合检测与预测框架。传统方法将检测与预测分步处理,难以利用时序信息。ForeSight通过多任务流式双向学习,使检测与预测共享查询记忆并无缝传递信息。预测感知的检测变换器通过多假设预测队列增强空间推理,而流式预测变换器则利用历史预测和优化后的检测提升时序一致性。与基于跟踪的方法不同,ForeSight采用无跟踪模型,消除显式目标关联,减少误差传播,并高效扩展至多帧序列。在nuScenes数据集上的实验表明,ForeSight达到54.9% EPA,超越此前方法9.3个百分点,同时在多视角检测与预测模型中取得最佳mAP和minADE表现。

原文摘要 · Abstract (English)

We introduce ForeSight, a novel joint detection and forecasting framework for vision-based 3D perception in autonomous vehicles. Traditional approaches treat detection and forecasting as separate sequential tasks, limiting their ability to leverage temporal cues. ForeSight addresses this limitation with a multi-task streaming and bidirectional learning approach, allowing detection and forecasting to share query memory and propagate information seamlessly. The forecast-aware detection transformer enhances spatial reasoning by integrating trajectory predictions from a multiple hypothesis forecast memory queue, while the streaming forecast transformer improves temporal consistency using past forecasts and refined detections. Unlike tracking-based methods, ForeSight eliminates the need for explicit object association, reducing error propagation with a tracking-free model that efficiently scales across multi-frame sequences. Experiments on the nuScenes dataset show that ForeSight achieves state-of-the-art performance, achieving an EPA of 54.9%, surpassing previous methods by 9.3%, while also attaining the best mAP and minADE among multi-view detection and forecasting models.

目标检测轨迹预测自动驾驶多视角感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。