arXiv:2602.07938cs.CVcs.RO2026-02

融合专用与通用预测,用动态占用栅格提升复杂场景预判能力

Integrating Specialized and Generic Agent Motion Prediction with Dynamic Occupancy Grid Maps

  • 基于动态占用栅格构建统一框架,同步预测车辆与场景未来状态
  • 在nuScenes和Woven Planet数据集上,对车辆与动态元素预测精度领先基准方法
  • 适合自动驾驶系统开发人员,尤其关注多智能体行为建模的团队

由于传感器数据不确定性、智能体行为复杂性及多种可行未来路径的存在,准确预测驾驶场景极具挑战。现有基于占用栅格的方法多聚焦于无差别场景预测,而专用预测虽能提供行为细节但难以泛化至感知差或未识别的智能体。为此,本文提出一种统一框架,通过轻量级时空主干网络,在简化的时间解码流程中同时预测未来占用状态栅格、车辆栅格与场景流栅格。设计了一种定制化的相互依赖损失函数,捕捉栅格间关联并生成多样化未来预测。利用占用状态信息引导流场演化,该损失函数作为正则项,确保占据变化符合障碍物与遮挡约束。模型不仅可预测车辆特定行为,还能识别其他动态实体并推演其演变。在真实世界nuScenes和Woven Planet数据集上的评估表明,本方法在动态车辆及通用动态场景元素预测上均优于基线方法。

原文摘要 · Abstract (English)

Accurate prediction of driving scene is a challenging task due to uncertainty in sensor data, the complex behaviors of agents, and the possibility of multiple feasible futures. Existing prediction methods using occupancy grid maps primarily focus on agent-agnostic scene predictions, while agent-specific predictions provide specialized behavior insights with the help of semantic information. However, both paradigms face distinct limitations: agent-agnostic models struggle to capture the behavioral complexities of dynamic actors, whereas agent-specific approaches fail to generalize to poorly perceived or unrecognized agents; combining both enables robust and safer motion forecasting. To address this, we propose a unified framework by leveraging Dynamic Occupancy Grid Maps within a streamlined temporal decoding pipeline to simultaneously predict future occupancy state grids, vehicle grids, and scene flow grids. Relying on a lightweight spatiotemporal backbone, our approach is centered on a tailored, interdependent loss function that captures inter-grid dependencies and enables diverse future predictions. By using occupancy state information to enforce flow-guided transitions, the loss function acts as a regularizer that directs occupancy evolution while accounting for obstacles and occlusions. Consequently, the model not only predicts the specific behaviors of vehicle agents, but also identifies other dynamic entities and anticipates their evolution within the complex scene. Evaluations on real-world nuScenes and Woven Planet datasets demonstrate superior prediction performances for dynamic vehicles and generic dynamic scene elements compared to baseline methods.

运动预测占用栅格自动驾驶多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。