让机器人从看懂到行动,通过空间推理提升操作泛化能力
From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation
- 通过空间关系推理生成中间表征,指导机器人精细操作
- 在8个基准上表现优异,真实场景任务成功率72%
- 零样本迁移效果显著,比最强基线提升30%
实现机器人操作的泛化仍是关键挑战,尤其在未见过的场景和新任务中。现有视觉-语言-动作(VLA)模型虽基于通用视觉-语言模型(VLM),但因具身数据集稀缺且异质性高,难以实现鲁棒的零样本性能。为此,我们提出FSD(From Seeing to Doing)——一种新型视觉-语言模型,通过空间关系推理生成中间表征,为机器人操作提供细粒度指导。方法结合分层数据训练流程与自一致性机制,将空间坐标与视觉信号对齐。大量实验验证了FSD在“看”与“做”两方面的综合能力,在8个通用空间推理与具身指代基准上表现卓越,并在我们提出的更复杂基准VABench上取得突破。此外,零样本机器人操作验证显示,其在SimplerEnv中达到40.6%成功率,在8个真实任务中平均达72%,显著优于最强基线,提升30%。
原文摘要 · Abstract (English)
Achieving generalization in robotic manipulation remains a critical challenge, particularly for unseen scenarios and novel tasks. Current Vision-Language-Action (VLA) models, while building on top of general Vision-Language Models (VLMs), still fall short of achieving robust zero-shot performance due to the scarcity and heterogeneity prevalent in embodied datasets. To address these limitations, we propose FSD (From Seeing to Doing), a novel vision-language model that generates intermediate representations through spatial relationship reasoning, providing fine-grained guidance for robotic manipulation. Our approach combines a hierarchical data pipeline for training with a self-consistency mechanism that aligns spatial coordinates with visual signals. Through extensive experiments, we comprehensively validated FSD's capabilities in both "seeing" and "doing," achieving outstanding performance across 8 benchmarks for general spatial reasoning and embodied reference abilities, as well as on our proposed more challenging benchmark VABench. We also verified zero-shot capabilities in robot manipulation, demonstrating significant performance improvements over baseline methods in both SimplerEnv and real robot settings. Experimental results show that FSD achieves 40.6% success rate in SimplerEnv and 72% success rate across 8 real-world tasks, outperforming the strongest baseline by 30%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。