arXiv:2504.20525cs.CV2025-04被引 1

利用视频序列时间信息突破单目3D车道检测模糊性。

Breaking Down Monocular Ambiguity: Exploiting Temporal Evolution for 3D Lane Detection

  • 通过多帧几何一致性建模恢复深度,构建可靠3D场景表示。
  • 融合历史与当前帧实例线索,显著提升远距离车道完整性。
  • 创新伪未来视角生成机制,有效捕捉易遗漏车道。

单目3D车道检测旨在从前视图像中估计车道的3D位置。然而,现有方法受限于单帧输入的固有模糊性,导致几何预测不准确、车道完整性差,尤其在远距离时更为明显。为此,我们提出利用车辆行驶过程中场景的时间演化信息来突破限制。所提出的几何感知时序聚合网络(GTA-Net)系统性地挖掘互补视角下的时序信息:首先,时序几何增强模块(TGEM)学习连续帧间的几何一致性,通过运动恢复深度,构建可靠的3D场景表示;其次,为增强车道完整性,时序实例感知查询生成模块(TIQG)聚合过去与当前帧的实例线索。关键在于,针对当前视图中模糊的车道,TIQG创新性地合成伪未来视角,生成可揭示被遮挡或缺失车道的查询。实验表明,GTA-Net在多个基准上达到新SOTA,显著优于现有单目3D车道检测方法。

原文摘要 · Abstract (English)

Monocular 3D lane detection aims to estimate the 3D position of lanes from frontal-view (FV) images. However, existing methods are fundamentally constrained by the inherent ambiguity of single-frame input, which leads to inaccurate geometric predictions and poor lane integrity, especially for distant lanes. To overcome this, we propose to unlock the rich information embedded in the temporal evolution of the scene as the vehicle moves. Our proposed Geometry-aware Temporal Aggregation Network (GTA-Net) systematically leverages the temporal information from complementary perspectives. First, Temporal Geometry Enhancement Module (TGEM) learns geometric consistency across consecutive frames, effectively recovering depth information from motion to build a reliable 3D scene representation. Second, to enhance lane integrity, Temporal Instance-aware Query Generation (TIQG) module aggregates instance cues from past and present frames. Crucially, for lanes that are ambiguous in the current view, TIQG innovatively synthesizes a pseudo future perspective to generate queries that reveal lanes which would otherwise be missed. The experiments demonstrate that GTA-Net achieves new SoTA results, significantly outperforming existing monocular 3D lane detection solutions.

3D车道检测单目视觉时序建模自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。