arXiv:2508.08697cs.CV2025-08ICRA被引 6

仅用RGB图像实现高速高精度非铺装路面空域检测

ROD: RGB-Only Fast and Efficient Off-road Freespace Detection

  • 基于预训练ViT提取RGB图像特征,结合轻量解码器提升效率
  • 在ORFD和RELLIS-3D上达到新SOTA,推理速度达50 FPS
  • 适合对实时性要求高的户外机器人导航场景

非铺装路面空域检测比铺装路面更具挑战性,因可通行区域边界模糊。以往先进方法依赖RGB图像与激光雷达(LiDAR)的多模态融合,但计算表面法向图导致推理延迟显著增加,难以满足实时应用需求,尤其在需要更高帧率的野外场景中。本文提出一种全新的仅使用RGB图像的非铺装路面空域检测方法ROD,摆脱对激光雷达及其计算开销的依赖。具体而言,采用预训练视觉变换器(ViT)从RGB图像中提取丰富特征,并设计一个轻量高效的解码器,在保证精度的同时大幅提升推理速度。ROD在ORFD和RELLIS-3D数据集上均达到新最优性能,推理速度达50 FPS,显著优于先前模型。

原文摘要 · Abstract (English)

Off-road freespace detection is more challenging than on-road scenarios because of the blurred boundaries of traversable areas. Previous state-of-the-art (SOTA) methods employ multi-modal fusion of RGB images and LiDAR data. However, due to the significant increase in inference time when calculating surface normal maps from LiDAR data, multi-modal methods are not suitable for real-time applications, particularly in real-world scenarios where higher FPS is required compared to slow navigation. This paper presents a novel RGB-only approach for off-road freespace detection, named ROD, eliminating the reliance on LiDAR data and its computational demands. Specifically, we utilize a pre-trained Vision Transformer (ViT) to extract rich features from RGB images. Additionally, we design a lightweight yet efficient decoder, which together improve both precision and inference speed. ROD establishes a new SOTA on ORFD and RELLIS-3D datasets, as well as an inference speed of 50 FPS, significantly outperforming prior models.

语义分割自动驾驶视觉感知实时检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。