arXiv:2606.08666cs.RO2026-06

用自然语言生成3D空间信念分布,提升机器人对远距物体定位的准确性。

Language as a Sensor: Calibrated Spatial Belief Estimation in 3D Scenes from Natural Language

论文配图:Language as a Sensor: Calibrated Spatial Belief Estimation in 3D Scenes from Natural Language
图 1 · 摘自论文原文
  • 将语言描述转化为带不确定性的空间概率分布,量化指代模糊与位置误差。
  • 在真实机器人上验证,融合语言后目标定位概率提升70%以上。
  • 适合需要理解人类口语指令的自主导航、人机协作场景。

部署于人机共处环境中的机器人常接收超出感知范围的自然语言空间描述(如“我把背包放在桌子上”)。传统度量语义地图忽略此类信息,而现有多模态模型在3D空间推理能力有限,且难以与其它传感器融合。为此,我们提出语言传感器模型(LSM),将每条话语及其场景图上下文映射为多模态分布,混合权重表示指代歧义(如“哪张桌子”),分量协方差编码空间不确定性(如“在桌上”的具体位置)。进一步提出VL-Map(视觉-语言度量语义地图)框架,将语言预测视为随机观测,与车载感知数据在统一信念图中融合。在VLA-3D基准及真实移动机器人上,LSM是唯一保持校准协方差估计的模型;融入VL-Map后,目标位置预测准确率显著提升,真值目标上的概率质量比最强基线模型高出约70%。

原文摘要 · Abstract (English)

Robots deployed in human-centric environments routinely receive natural-language descriptions of spatial information ("I left my backpack on the table") that reference parts of the world beyond their perceptual field of view. Traditional metric-semantic mapping ignores this signal, while off-the-shelf multimodal models remain limited in 3D spatial reasoning and are not directly amenable to fusion with other sensor modalities. To convert language observations into a calibrated spatial distribution, we train a Language Sensor Model (LSM) that maps each utterance and its scene-graph context to a multimodal distribution, with mixture weights encoding referential ambiguity (e.g., "which table") and component covariances encoding spatial uncertainty (e.g., where "on the table" the target lies). We then introduce VL-Map (Vision-Language Metric-Semantic Mapping), a probabilistic framework that treats these language predictions as stochastic observations and fuses them with onboard perception within a unified belief map. On the VLA-3D benchmark as well as on a real-world mobile robot, LSM is the only language predictor whose covariance estimates remain within the calibrated regime; fused into VL-Map, it leads to more accurate predictions of the target object location (~70% more probability mass on the true target compared to the strongest foundation-model baseline).

3D空间理解语言感知机器人定位多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。