用视觉语言模型让助老助残机器人读懂房间语义,自动导航不迷路。
Open-Vocabulary Semantic Segmentation with Uncertainty Alignment for Robotic Scene Understanding in Indoor Building Environments
- 基于视觉语言模型与大语言模型,实现开放词汇的场景分割与识别
- 提出'分段-检测-选择'框架,提升对模糊场景的理解能力
- 特别适合需要理解口语指令的智能轮椅等辅助机器人使用
随着残障人士数量上升,对先进辅助技术的需求日益增长。自主助行机器人如智能轮椅需具备强大的空间分割与语义识别能力,以在复杂建筑环境中有效导航。场景分割需划分出房间或功能区,语义识别则为这些区域赋予标签,从而实现针对用户需求的精准定位。现有方法多依赖深度学习,但封闭词汇系统难以理解自然口语指令,且普遍忽略场景识别中的不确定性,导致在复杂环境成功率低。为此,本文提出一种基于视觉语言模型(VLMs)与大语言模型(LLMs)的开放词汇场景语义分割与检测流水线。采用'分段-检测-选择'框架,支持对非预设词汇的动态理解,提升机器人在建筑环境中的自适应与直观导航能力。
原文摘要 · Abstract (English)
The global rise in the number of people with physical disabilities, in part due to improvements in post-trauma survivorship and longevity, has amplified the demand for advanced assistive technologies to improve mobility and independence. Autonomous assistive robots, such as smart wheelchairs, require robust capabilities in spatial segmentation and semantic recognition to navigate complex built environments effectively. Place segmentation involves delineating spatial regions like rooms or functional areas, while semantic recognition assigns semantic labels to these regions, enabling accurate localization to user-specific needs. Existing approaches often utilize deep learning; however, these close-vocabulary detection systems struggle to interpret intuitive and casual human instructions. Additionally, most existing methods ignore the uncertainty of the scene recognition problem, leading to low success rates, particularly in ambiguous and complex environments. To address these challenges, we propose an open-vocabulary scene semantic segmentation and detection pipeline leveraging Vision Language Models (VLMs) and Large Language Models (LLMs). Our approach follows a 'Segment Detect Select' framework for open-vocabulary scene classification, enabling adaptive and intuitive navigation for assistive robots in built environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。