arXiv:2607.16173cs.RO2026-07被引 1

让3D地图能理解物体如何动,支持语言查询运动状态。

Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps

论文配图:Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps
图 1 · 摘自论文原文
  • 融合语义先验与实际运动观测,生成可查询的动态属性
  • 在仿真和真实数据上均提升运动判断准确率,减少误报
  • 适合需要感知动态变化的机器人导航任务

开放词汇的3D地图能让机器人回答关于物体位置和身份的语言问题,但假设世界静态,无法回答行为动态问题。我们提出视觉-语言-运动地图(VLMM),一种开放词汇、语言可查询的3D地图,通过规则驱动的意图路由系统实现对开放词汇物体名词的查询,而非通用自然语言接口。每个元素携带融合的运动属性:由视觉-语言模型(VLM)/大语言模型(LLM)提供的语义可动性先验,结合几何观测到的跨帧运动,以及每元素的不确定性。查询转化为属性过滤器,区分已观察移动、可能移动但未动、保持静止的物体。在带精确真值的控制模拟器基准(AI2-THOR,三种场景类型)上,消融实验表明各字段不可替代:仅依赖语义的基线即使使用强特征也无法完成运动查询;语义先验无法回答“什么在动”,观测运动也无法回答“什么可能动”。在真实动态RGB-D数据(TUM和Bonn,六段序列)上,不确定性通道显著提升动/静分类平均精度,减少误判,并对估计噪声姿态具有鲁棒性。原始置信度未经校准,但事后等距校准后达到0.10的期望校准误差。VLMM是表示层面的贡献:现有最接近的映射至少缺少四个特性之一——开放词汇、语言可查询、融合先验与观测运动、每元素不确定性——而我们的组合完整提供。

原文摘要 · Abstract (English)

Open-vocabulary 3D maps let robots answer language queries about what and where, but they assume a static world and cannot answer queries about how scene elements behave. We introduce Vision-Language-Motion Maps (VLMM), an open-vocabulary, language-queryable 3D map - queried through a rule-based intent router over open-vocabulary object nouns, not a general natural-language interface - in which each element carries a fused motion attribute: a VLM/LLM semantic movability prior combined with geometrically observed cross-frame motion, together with a per-element uncertainty. Queries reduce to attribute filters that distinguish what has been seen to move, what could move but has not, and what stays still. On a controlled simulator benchmark with exact ground truth (AI2-THOR, three scene types) we show through ablation that the schema fields are non-substitutable: a semantic-only baseline fails motion queries even with strong features, and neither motion field substitutes for the other (the prior cannot answer "what is moving," observed motion cannot answer "what could move"). On real dynamic RGB-D (TUM and Bonn, six sequences) we show the uncertainty channel - our key difference from prior fused-motion work - consistently improves moving-vs-static average precision and reduces false motion flags, and that it is robust to estimated (noisy) poses. The raw confidence is not calibrated, but post-hoc isotonic calibration reaches an expected calibration error of 0.10. VLMM is a representation contribution: the closest prior maps each lack at least one of the four properties - open-vocabulary, language-queryable, fused prior-and-observed motion, and per-element uncertainty - that our combination provides.

3D地图运动感知机器人不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。