提出统一框架,让车道与交通标志推理更连贯准确。
Geometry-Guided Representations for Coherent Lane and Traffic Topology Reasoning in Driving Scenes
- 用几何引导注意力增强车道特征表达
- 在多个指标上提升3.1至5.3分,效果显著
- 适合自动驾驶中需要精准拓扑理解的场景
道路拓扑推理对自动驾驶至关重要,需同时精准感知道路元素并理解其复杂连接关系,包括车道间连接(L2L)和车道与交通标志的关联(L2T)。现有方法常将感知与拓扑推理割裂,忽视二者相互促进潜力。尤其在特征提取阶段,多数工作忽略几何关系,依赖脆弱的后处理或仅在推理时应用坐标启发式规则。为此,本文提出统一框架CoPo(Coherent Perception and toPology),从三个层次实现几何引导的关联建模:1)感知层,引入关系感知车道检测器,通过几何偏置自注意力和曲线引导交叉注意力,融合结构先验;2)推理层,设计增强的关系拓扑头,包括几何增强的L2L头和跨视角L2T头,有效对齐特征以推断连接性;3)监督层,采用对比信息熵(InfoNCE)策略正则化关系嵌入,使相连样本在隐空间更接近。该协同多层级设计支持端到端联合优化。在OpenLane-V2上的大量实验表明,CoPo显著优于现有方法,分别取得DET$_l$ +3.1、TOP$_{ll}$ +5.3、TOP$_{lt}$ +4.9、OLS +4.4的提升,刷新当前最佳性能。
原文摘要 · Abstract (English)
Road topology reasoning is fundamental for autonomous driving, requiring both accurate perception of road elements and understanding of their complex connectivity, including lane connectivity (Lane-to-Lane, L2L) and traffic regulation (Lane-to-Traffic signs, L2T). However, existing methods typically treat perception and topology reasoning as fragmented tasks, ignoring their potential for mutual enhancement. Crucially, while topology is inherently relational, prior works often overlook geometric relationships during feature extraction, relying instead on brittle post-processing or coordinate-based heuristics applied only at inference time. To bridge this gap, we propose CoPo (Coherent Perception and toPology), a unified framework that integrates geometry-guided relational modeling across three levels: 1) Perception-level: We introduce a relation-aware lane detector that utilizes geometry-biased self-attention and curve-guided cross-attention to enrich lane representations with structural priors; 2) Reasoning-level: We design relation-enhanced topology heads, including a geometry-enhanced L2L head and a cross-view L2T head, which effectively align features to infer connectivity; and 3) Supervision-level: We implement a contrastive InfoNCE strategy to regularize relational embeddings, pulling connected pairs closer in the latent space. This coherent multi-level design enables end-to-end joint optimization of perception and reasoning. Extensive experiments on OpenLane-V2 demonstrate that CoPo significantly outperforms existing methods, achieving gains of {\textbf{+3.1}} in DET$_l$, {\textbf{+5.3}} in TOP$_{ll}$, {\textbf{+4.9}} in TOP$_{lt}$, and {\textbf{+4.4}} overall in OLS, setting a new state-of-the-art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。