让机器人理解人类语言的空间不确定性,提升协作感知精度。
Spatial Language Likelihood Grounding Network for Bayesian Fusion of Human-Robot Observations
- 通过特征金字塔网络学习语言与地图特征的关联,生成空间语义的置信度
- 在均值NLL上达到专家规则水平,标准差更低,鲁棒性更强
- 适合需要融合人类描述与机器人传感器数据的协作场景
融合人类观察信息可帮助机器人克服协作任务中的感知局限。然而,具备不确定性感知的融合框架需要一个基于语义的置信度表示来刻画人类输入的不确定性。本文提出特征金字塔似然建模网络(FP-LGN),通过学习地图图像特征及其与空间关系语义的关系,实现空间语言的语义接地。模型采用三阶段课程学习训练为概率估计器,捕捉人类语言中的随机不确定性。实验表明,FP-LGN在平均负对数似然(NLL)上达到专家设计规则水平,并展现出更低的标准差,更具鲁棒性。协作感知结果表明,该接地似然成功实现了异构人类语言与机器人传感器数据的不确定性感知融合,显著提升了人-机协作任务性能。
原文摘要 · Abstract (English)
Fusing information from human observations can help robots overcome sensing limitations in collaborative tasks. However, an uncertainty-aware fusion framework requires a grounded likelihood representing the uncertainty of human inputs. This paper presents a Feature Pyramid Likelihood Grounding Network (FP-LGN) that grounds spatial language by learning relevant map image features and their relationships with spatial relation semantics. The model is trained as a probability estimator to capture aleatoric uncertainty in human language using three-stage curriculum learning. Results showed that FP-LGN matched expert-designed rules in mean Negative Log-Likelihood (NLL) and demonstrated greater robustness with lower standard deviation. Collaborative sensing results demonstrated that the grounded likelihood successfully enabled uncertainty-aware fusion of heterogeneous human language observations and robot sensor measurements, achieving significant improvements in human-robot collaborative task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。