arXiv:2509.19378cs.CVcs.AR2025-09

用深度学习让无人车在矿山等复杂野外环境实时识别可行驶区域。

Vision-Based Perception for Autonomous Vehicles in Off-Road Environment Using Deep Learning

  • 提出模块化网络框架CMSNet,灵活适配不同地形感知需求。
  • 在近1.2万张含雨夜尘土的图像上实现高精度实时语义分割。
  • 适用于矿区、发展中国家等无预设路径的恶劣野外场景。

为应对露天矿及发展中国家非铺装路面的自动驾驶需求,本文提出一种面向非铺装与野外环境的视觉感知系统,可在无预设路径条件下导航复杂地形。提出可配置模块化分割网络(CMSNet)框架,支持多种结构配置。在包含近12,000张图像的新数据集Kamino上,对夜间、雨天、扬尘等恶劣条件下的新图像进行障碍物与可通行区域分割训练。研究了无明确道路边界时的可行驶区域检测算法行为,评估了真实环境下的实时语义分割性能。通过使用TensorRT、C++和CUDA对CMSNet的卷积层进行系统性优化,实现了实时推理。在两个数据集上的实证实验验证了系统的有效性。

原文摘要 · Abstract (English)

Low-latency intelligent systems are required for autonomous driving on non-uniform terrain in open-pit mines and developing countries. This work proposes a perception system for autonomous vehicles on unpaved roads and off-road environments, capable of navigating rough terrain without a predefined trail. The Configurable Modular Segmentation Network (CMSNet) framework is proposed, facilitating different architectural arrangements. CMSNet configurations were trained to segment obstacles and trafficable ground on new images from unpaved/off-road scenarios with adverse conditions (night, rain, dust). We investigated applying deep learning to detect drivable regions without explicit track boundaries, studied algorithm behavior under visibility impairment, and evaluated field tests with real-time semantic segmentation. A new dataset, Kamino, is presented with almost 12,000 images from an operating vehicle with eight synchronized cameras. The Kamino dataset has a high number of labeled pixels compared to similar public collections and includes images from an off-road proving ground emulating a mine under adverse visibility. To achieve real-time inference, CMSNet CNN layers were methodically removed and fused using TensorRT, C++, and CUDA. Empirical experiments on two datasets validated the proposed system's effectiveness.

自动驾驶视觉感知野外环境语义分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。