用隐式世界模型提升自动驾驶车道拓扑推理的时序感知能力
FASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World Models
- 分快慢两路并行处理,融合历史与新初始化查询实现互补监督
- 引入动作条件的隐式世界模型,显著提升时序状态传播性能
- 在OpenLane-V2上车道段检测和中心线感知均超越现有方法
车道段拓扑推理可提供全面的鸟瞰视角道路场景理解,是面向规划的端到端自动驾驶系统的关键感知模块。现有方法在利用时序信息提升检测与推理性能方面仍存在不足。近期基于流的时序传播方法虽在查询与鸟瞰图层面引入时序线索取得进展,但仍受限于对历史查询的过度依赖、姿态估计失败的敏感性以及时序传播不充分。为此,本文提出FASTopoWM,一种基于隐式世界模型增强的快慢车道段拓扑推理框架。为降低姿态估计误差影响,该统一框架实现历史与新初始化查询的并行监督,促进快慢系统间相互强化。此外,设计了基于动作隐变量的隐式查询与鸟瞰图世界模型,将过去观测的状态表示传播至当前时刻。在OpenLane-V2基准上的大量实验表明,FASTopoWM在车道段检测(mAP 37.4% vs. 33.6%)和中心线感知(OLS 46.3% vs. 41.5%)上均优于现有最先进方法。
原文摘要 · Abstract (English)
Lane segment topology reasoning provides comprehensive bird's-eye view (BEV) road scene understanding, which can serve as a key perception module in planning-oriented end-to-end autonomous driving systems. Existing lane topology reasoning methods often fall short in effectively leveraging temporal information to enhance detection and reasoning performance. Recently, stream-based temporal propagation method has demonstrated promising results by incorporating temporal cues at both the query and BEV levels. However, it remains limited by over-reliance on historical queries, vulnerability to pose estimation failures, and insufficient temporal propagation. To overcome these limitations, we propose FASTopoWM, a novel fast-slow lane segment topology reasoning framework augmented with latent world models. To reduce the impact of pose estimation failures, this unified framework enables parallel supervision of both historical and newly initialized queries, facilitating mutual reinforcement between the fast and slow systems. Furthermore, we introduce latent query and BEV world models conditioned on the action latent to propagate the state representations from past observations to the current timestep. This design substantially improves the performance of temporal perception within the slow pipeline. Extensive experiments on the OpenLane-V2 benchmark demonstrate that FASTopoWM outperforms state-of-the-art methods in both lane segment detection (37.4% v.s. 33.6% on mAP) and centerline perception (46.3% v.s. 41.5% on OLS).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。