综述语义视觉定位与建图最新进展及挑战
Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- 按模块化流程梳理语义SLAM技术架构
- 分析深度学习与大模型在其中的应用潜力
- 适合研究机器人环境感知的学者参考
语义同时定位与建图(Semantic SLAM)是机器人学与计算机视觉中的关键研究领域,致力于同步实现机器人定位,并为环境构建包含语义信息的完整模型。自二十多年前首个基础性工作出现以来,该领域受到越来越多科学界的关注。尽管意义重大,但现有综述尚未全面涵盖近期进展与持续挑战。本文系统梳理了语义SLAM的最新技术发展,旨在揭示当前趋势与核心障碍。从视觉SLAM演进出发,分析其优势与特点,同时批判性评估已有综述。提出统一的问题建模与模块化解决方案框架,将任务分解为视觉定位、语义特征提取、建图、数据关联和回环优化等阶段。此外,探讨深度学习与大型语言模型等替代方法,并综述主流SLAM数据集的研究进展。最后讨论未来可能的研究方向,为希望深入该领域的研究人员提供全面参考。
原文摘要 · Abstract (English)
Semantic Simultaneous Localization and Mapping (SLAM) is a critical area of research within robotics and computer vision, focusing on the simultaneous localization of robotic systems and associating semantic information to construct the most accurate and complete comprehensive model of the surrounding environment. Since the first foundational work in Semantic SLAM appeared more than two decades ago, this field has received increasing attention across various scientific communities. Despite its significance, the field lacks comprehensive surveys encompassing recent advances and persistent challenges. In response, this study provides a thorough examination of the state-of-the-art of Semantic SLAM techniques, with the aim of illuminating current trends and key obstacles. Beginning with an in-depth exploration of the evolution of visual SLAM, this study outlines its strengths and unique characteristics, while also critically assessing previous survey literature. Subsequently, a unified problem formulation and evaluation of the modular solution framework is proposed, which divides the problem into discrete stages, including visual localization, semantic feature extraction, mapping, data association, and loop closure optimization. Moreover, this study investigates alternative methodologies such as deep learning and the utilization of large language models, alongside a review of relevant research about contemporary SLAM datasets. Concluding with a discussion on potential future research directions, this study serves as a comprehensive resource for researchers seeking to navigate the complex landscape of Semantic SLAM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。