用单目视频实现动态环境中的开放词汇语义定位,无需标定或深度传感器。
RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments

- 基于视觉语言基础模型,实时融合多模态信息与几何结构。
- 在动态场景中保持地图一致性,在TUM-RGBD上达最新水平。
- 适合无人机器人、野外视频分析等真实场景应用。
我们提出RADIO-ViPE(Reduce All Domains Into One -- Video Pose Engine),一种在线语义SLAM系统,可在动态环境中实现几何感知的开放词汇定位,将任意自然语言查询与3D空间中的区域和物体关联。不同于需校准的RGB-D输入方法,RADIO-ViPE直接处理原始单目RGB视频流,无需相机内参、深度传感器或位姿初始化。系统将来自聚合基础模型(如RADIO)的视觉-语言嵌入与几何场景信息紧密耦合,贯穿初始化、优化及因子图连接过程,提升多模态地图一致性。优化过程采用自适应鲁棒核,有效应对主动移动物体及代理位移的场景元素(如会话期间家具重排)。实验表明,RADIO-ViPE在动态TUM-RGBD基准上达到当前最优性能,且优于依赖校准数据与静态假设的离线方法。该系统填补了实际部署的关键空白,为自主机器人和无约束野外视频流提供鲁棒的开放词汇语义定位能力。
原文摘要 · Abstract (English)
We present RADIO-ViPE (Reduce All Domains Into One -- Video Pose Engine), an online semantic SLAM system that enables geometry-aware open-vocabulary grounding, associating arbitrary natural language queries with localized 3D regions and objects in dynamic environments. Unlike existing approaches that require calibrated, posed RGB-D input, RADIO-ViPE operates directly on raw monocular RGB video streams, requiring no prior camera intrinsics, depth sensors, or pose initialization. The system tightly couples multi-modal embeddings -- spanning vision and language -- derived from agglomerative foundation models (e.g., RADIO) with geometric scene information. This coupling takes place in initialization, optimization and factor graph connections to improve the consistency of the map from multiple modalities. The optimization is wrapped within adaptive robust kernels, designed to handle both actively moving objects and agent-displaced scene elements (e.g., furniture rearranged during ego-centric session). Experiments demonstrate that RADIO-ViPE achieves state-of-the-art results on the dynamic TUM-RGBD benchmark while maintaining competitive performance against offline open-vocabulary methods that rely on calibrated data and static scene assumptions. RADIO-ViPE bridges a critical gap in real-world deployment, enabling robust open-vocabulary semantic grounding for autonomous robotics and unconstrained in-the-wild video streams. Project page: https://be2rlab.github.io/radio_vipe
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。