arXiv:2503.16711cs.ROcs.CV2025-03被引 1

用深度信息提升自动驾驶决策鲁棒性,实测效果显著。

Depth Matters: Multimodal RGB-D Perception for Robust Autonomous Agents

  • 早期融合RGB与深度图,构建轻量时序控制网络。
  • 在帧丢失和噪声干扰下仍保持稳定控制性能。
  • 适合部署于真实硬件的高可靠性自主系统。

依赖感知进行实时控制决策的自主智能体需要高效且鲁棒的架构。本文表明,将深度信息融入RGB输入能显著提升智能体预测转向指令的能力,优于仅使用RGB的情况。我们评估了基于融合RGB-D特征的轻量级循环控制器,在序列决策中的表现。为训练模型,我们通过专家驾驶的小型自动驾驶汽车采集高质量数据,涵盖不同难度的转向场景。模型成功部署于真实硬件,在分布外条件下自然规避了动态与静态障碍物。具体发现显示,早期融合深度数据可构建高度鲁棒的控制器,即使在帧丢失和噪声增加时仍保持有效,且不分散网络对任务的关注。

原文摘要 · Abstract (English)

Autonomous agents that rely purely on perception to make real-time control decisions require efficient and robust architectures. In this work, we demonstrate that augmenting RGB input with depth information significantly enhances our agents' ability to predict steering commands compared to using RGB alone. We benchmark lightweight recurrent controllers that leverage the fused RGB-D features for sequential decision-making. To train our models, we collect high-quality data using a small-scale autonomous car controlled by an expert driver via a physical steering wheel, capturing varying levels of steering difficulty. Our models were successfully deployed on real hardware and inherently avoided dynamic and static obstacles, under out-of-distribution conditions. Specifically, our findings reveal that the early fusion of depth data results in a highly robust controller, which remains effective even with frame drops and increased noise levels, without compromising the network's focus on the task.

多模态感知自动驾驶深度学习鲁棒控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。