arXiv:2504.02477cs.ROcs.CV2025-04中稿 · Information Fusion综述被引 89

综述多模态融合与视觉语言模型在机器人视觉中的应用进展

Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision

  • 按任务导向梳理编码器-解码器、注意力机制、图网络等融合方法
  • 对比大模型驱动的VLM与传统融合方法在定位、导航等任务表现
  • 适合关注机器人感知与交互研究的学者参考

机器人视觉受益于多模态融合技术与视觉语言模型(VLMs)的发展。本文从任务导向视角系统回顾了这些技术在机器人视觉中的应用与进展。针对语义场景理解任务,将融合方法分为编码器-解码器框架、基于注意力的架构和图神经网络。同时分析了这些策略在同步定位与地图构建(SLAM)、3D目标检测、导航与操作等关键任务中的结构特征与实际实现。对比了基于大语言模型(LLMs)的VLM与传统多模态融合方法的演进路径与适用性。此外,深入评估了常用数据集在真实机器人场景中的适用性与挑战。基于此,识别出当前研究的关键挑战:跨模态对齐、高效融合、实时部署与域适应。提出未来方向:自监督学习以构建鲁棒多模态表征,结构化空间记忆与环境建模以增强空间智能,以及集成对抗鲁棒性与人类反馈机制以实现伦理对齐的系统部署。通过全面综述、对比分析与前瞻讨论,为推进机器人视觉中的多模态感知与交互提供重要参考。完整文献列表见 https://github.com/Xiaofeng-Han-Res/MF-RV。

原文摘要 · Abstract (English)

Robot vision has greatly benefited from advancements in multimodal fusion techniques and vision-language models (VLMs). We adopt a task-oriented perspective to systematically review the applications and advancements of multimodal fusion methods and VLMs in the field of robot vision. For semantic scene understanding tasks, we categorize fusion approaches into encoder-decoder frameworks, attention-based architectures, and graph neural networks. Meanwhile, we also analyze the architectural characteristics and practical implementations of these fusion strategies in key tasks such as simultaneous localization and mapping (SLAM), 3D object detection, navigation, and manipulation. We compare the evolutionary paths and applicability of VLMs based on large language models (LLMs) with traditional multimodal fusion methods.Additionally, we conduct an in-depth analysis of commonly used datasets, evaluating their applicability and challenges in real-world robotic scenarios. Building on this analysis, we identify key challenges in current research, including cross-modal alignment, efficient fusion, real-time deployment, and domain adaptation. We propose future directions such as self-supervised learning for robust multimodal representations, structured spatial memory and environment modeling to enhance spatial intelligence, and the integration of adversarial robustness and human feedback mechanisms to enable ethically aligned system deployment. Through a comprehensive review, comparative analysis, and forward-looking discussion, we provide a valuable reference for advancing multimodal perception and interaction in robotic vision. A comprehensive list of studies in this survey is available at https://github.com/Xiaofeng-Han-Res/MF-RV.

机器人视觉多模态融合视觉语言模型SLAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。