无需深度传感器,实时融合视觉、语言与几何信息实现动态环境语义定位
KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM
- 通过自适应鲁棒核将DINO特征与几何约束紧密耦合
- 在未标定单目视频流上实现在线定位与开放词汇语义建图
- 适合自主机器人与AR/VR,支持互联网规模训练
我们提出KM-ViPE(知识映射视频位姿引擎),一种适用于动态环境中未标定单目相机的实时开放词汇SLAM框架。不同于依赖深度传感器和离线标定的系统,KM-ViPE直接处理原始RGB流,特别适合第一人称应用,并可利用互联网规模视频数据进行训练。该系统通过基于高级特征的自适应鲁棒核,将DINO视觉特征与几何约束紧密耦合,有效处理移动物体及可移动静态物体(如第一人称视角中的移动家具)。通过融合几何与深度视觉特征并对其对齐语言嵌入,实现同步在线定位与开放词汇语义建图。实验结果与最先进方法相当,而现有方案或需离线运行、依赖深度数据或里程计估计,或缺乏对动态场景的鲁棒性。KM-ViPE受益于互联网规模训练,独特结合在线运行、未标定单目输入与动态场景鲁棒性,适用于自主机器人与AR/VR应用,推动具身AI的实际空间智能发展。
原文摘要 · Abstract (English)
We present KM-ViPE (Knowledge Mapping Video Pose Engine), a real-time open-vocabulary SLAM framework for uncalibrated monocular cameras in dynamic environments. Unlike systems requiring depth sensors and offline calibration, KM-ViPE operates directly on raw RGB streams, making it ideal for ego-centric applications and harvesting internet-scale video data for training. KM-ViPE tightly couples DINO visual features with geometric constraints through a high-level features based adaptive robust kernel that handles both moving objects and movable static objects (e.g., moving furniture in ego-centric views). The system performs simultaneous online localization and open-vocabulary semantic mapping by fusing geometric and deep visual features aligned with language embeddings. Our results are competitive with state-of-the-art approaches, while existing solutions either operate offline, need depth data and/or odometry estimation, or lack dynamic scene robustness. KM-ViPE benefits from internet-scale training and uniquely combines online operation, uncalibrated monocular input, and robust handling of dynamic scenes, which makes it a good fit for autonomous robotics and AR/VR applications and advances practical spatial intelligence capabilities for embodied AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。