arXiv:2409.12716cs.CVcs.AI2024-09被引 2

用单目相机的光流信息提升自动驾驶转向预测精度

Optical Flow Matters: an Empirical Comparative Study on Fusing Monocular Extracted Modalities for Better Steering

  • 融合图像与光流数据,采用早融合与混合融合策略
  • 在波士顿驾驶数据集上误差降低31%,优于无光流的方法
  • 适合关注视觉感知与决策融合的自动驾驶研究者

自动驾驶导航是人工智能的关键挑战,需具备鲁棒且准确的决策能力。本文提出一种新型端到端方法,仅通过单目相机提取多模态信息,提升自驾车的转向预测性能。不同于依赖多个传感器(成本高、结构复杂)或仅依赖RGB图像(在不同条件下鲁棒性不足)的传统模型,本方法显著提升了单一视觉传感器下的转向预测表现。通过融合RGB图像与深度补全信息或光流数据,构建了包含早期融合与混合融合的综合框架。采用三种神经网络架构实现:卷积神经网络-中性电路策略(CNN-NCP)、变分自编码器-长短期记忆(VAE-LSTM)和神经电路策略架构VAE-NCP。实验基于波士顿驾驶数据集的对比研究显示,融合图像与运动信息的模型具有强鲁棒性与可靠性,相比不使用光流的先进方法,转向估计误差降低31%。结果表明,结合先进神经网络结构(以CNN融合数据、递归网络从隐空间推断指令),光流数据可有效提升自动驾驶转向估计性能。

原文摘要 · Abstract (English)

Autonomous vehicle navigation is a key challenge in artificial intelligence, requiring robust and accurate decision-making processes. This research introduces a new end-to-end method that exploits multimodal information from a single monocular camera to improve the steering predictions for self-driving cars. Unlike conventional models that require several sensors which can be costly and complex or rely exclusively on RGB images that may not be robust enough under different conditions, our model significantly improves vehicle steering prediction performance from a single visual sensor. By focusing on the fusion of RGB imagery with depth completion information or optical flow data, we propose a comprehensive framework that integrates these modalities through both early and hybrid fusion techniques. We use three distinct neural network models to implement our approach: Convolution Neural Network - Neutral Circuit Policy (CNN-NCP) , Variational Auto Encoder - Long Short-Term Memory (VAE-LSTM) , and Neural Circuit Policy architecture VAE-NCP. By incorporating optical flow into the decision-making process, our method significantly advances autonomous navigation. Empirical results from our comparative study using Boston driving data show that our model, which integrates image and motion information, is robust and reliable. It outperforms state-of-the-art approaches that do not use optical flow, reducing the steering estimation error by 31%. This demonstrates the potential of optical flow data, combined with advanced neural network architectures (a CNN-based structure for fusing data and a Recurrence-based network for inferring a command from latent space), to enhance the performance of autonomous vehicles steering estimation.

自动驾驶光流多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。