arXiv:2504.19002cs.LGcs.CV2025-04被引 8

融合视觉与激光数据,提升机器人在复杂环境中的导航精度。

Deep Learning-Based Multi-Modal Fusion for Robust Robot Perception and Navigation

  • 轻量级网络提取多模态特征,增强表征能力
  • 自适应加权融合策略使系统更鲁棒,定位精度提升2.2%
  • 引入时序建模,改善动态场景感知效果

本文提出一种基于深度学习的多模态融合架构,旨在提升自主导航机器人在复杂环境中的感知能力。通过创新的特征提取模块、自适应融合策略和时序建模机制,有效融合RGB图像与LiDAR数据。主要贡献包括:(a) 设计轻量级特征提取网络以增强特征表达;(b) 提出自适应加权跨模态融合策略,提升系统鲁棒性;(c) 引入时序信息建模,提高动态场景感知精度。在KITTI数据集上的实验表明,所提方法使导航与定位精度分别提升3.5%和2.2%,且保持实时性能。该工作为复杂环境下自主机器人导航提供了新方案。

原文摘要 · Abstract (English)

This paper introduces a novel deep learning-based multimodal fusion architecture aimed at enhancing the perception capabilities of autonomous navigation robots in complex environments. By utilizing innovative feature extraction modules, adaptive fusion strategies, and time-series modeling mechanisms, the system effectively integrates RGB images and LiDAR data. The key contributions of this work are as follows: a. the design of a lightweight feature extraction network to enhance feature representation; b. the development of an adaptive weighted cross-modal fusion strategy to improve system robustness; and c. the incorporation of time-series information modeling to boost dynamic scene perception accuracy. Experimental results on the KITTI dataset demonstrate that the proposed approach increases navigation and positioning accuracy by 3.5% and 2.2%, respectively, while maintaining real-time performance. This work provides a novel solution for autonomous robot navigation in complex environments.

多模态融合机器人导航深度学习LiDAR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。