让机器人通过语言理解三维世界,实现更智能的自主感知。
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- 用大模型+三维视觉融合,让机器人理解环境语义
- 支持零样本3D分割与语言引导操作,提升交互能力
- 适合研究智能机器人、多模态感知的学者与工程师
随着人工智能与机器人技术的快速发展,大语言模型(LLMs)与三维视觉的融合正成为提升机器人感知能力的变革性路径。该综述系统分析了当前主流方法、应用及挑战,聚焦下一代机器人感知技术。首先介绍LLMs与三维数据表示的基础原理,深入探讨机器人关键的三维感知技术。重点涵盖场景理解、文本生成3D内容、物体定位与具身智能体等前沿进展,如零样本3D分割、动态场景合成与语言引导操控。此外,讨论融合触觉、听觉与热成像的多模态大模型,增强环境理解与决策能力。为支持未来研究,整理了专用于3D-语言与视觉任务的基准数据集与评估指标。最后指出关键挑战与方向:自适应模型架构、跨模态对齐优化与实时处理能力,推动更智能、情境感知与自主的机器人感知系统发展。
原文摘要 · Abstract (English)
With the rapid advancement of artificial intelligence and robotics, the integration of Large Language Models (LLMs) with 3D vision is emerging as a transformative approach to enhancing robotic sensing technologies. This convergence enables machines to perceive, reason and interact with complex environments through natural language and spatial understanding, bridging the gap between linguistic intelligence and spatial perception. This review provides a comprehensive analysis of state-of-the-art methodologies, applications and challenges at the intersection of LLMs and 3D vision, with a focus on next-generation robotic sensing technologies. We first introduce the foundational principles of LLMs and 3D data representations, followed by an in-depth examination of 3D sensing technologies critical for robotics. The review then explores key advancements in scene understanding, text-to-3D generation, object grounding and embodied agents, highlighting cutting-edge techniques such as zero-shot 3D segmentation, dynamic scene synthesis and language-guided manipulation. Furthermore, we discuss multimodal LLMs that integrate 3D data with touch, auditory and thermal inputs, enhancing environmental comprehension and robotic decision-making. To support future research, we catalog benchmark datasets and evaluation metrics tailored for 3D-language and vision tasks. Finally, we identify key challenges and future research directions, including adaptive model architectures, enhanced cross-modal alignment and real-time processing capabilities, which pave the way for more intelligent, context-aware and autonomous robotic sensing systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。