用动态热图表示空间目标,提升机器人任务的鲁棒性与成功率
More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
- 将空间目标建模为可变的连续热图,而非固定点
- 在真实场景中实现82%任务成功率,比之前快50倍
- 适合需要精准空间推理的机器人操控与导航任务
许多语言引导的机器人系统将空间推理简化为离散点,对感知噪声和语义模糊敏感。为此,我们提出RoboMAP框架,将空间目标表示为连续且自适应的可用性热图。这种密集表示捕捉了空间定位中的不确定性,为下游策略提供更丰富的信息,显著提升了任务成功率和可解释性。RoboMAP在多数基准测试中超越现有最优方法,速度最高提升50倍,并在真实世界操作中达到82%的成功率。在大量模拟与物理实验中表现出稳健性能,并展现出强大的零样本泛化能力,适用于导航任务。更多信息与视频见https://robo-map.github.io。
原文摘要 · Abstract (English)
Many language-guided robotic systems rely on collapsing spatial reasoning into discrete points, making them brittle to perceptual noise and semantic ambiguity. To address this challenge, we propose RoboMAP, a framework that represents spatial targets as continuous, adaptive affordance heatmaps. This dense representation captures the uncertainty in spatial grounding and provides richer information for downstream policies, thereby significantly enhancing task success and interpretability. RoboMAP surpasses the previous state-of-the-art on a majority of grounding benchmarks with up to a 50x speed improvement, and achieves an 82\% success rate in real-world manipulation. Across extensive simulated and physical experiments, it demonstrates robust performance and shows strong zero-shot generalization to navigation. More details and videos can be found at https://robo-map.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。