让机器人像人一样一步步思考物体关系,提升复杂操作能力。
GSR: Learning Structured Reasoning for Embodied Manipulation
- 用场景图显式建模物体状态与空间关系,逐步推理动作条件
- 在4个数据集上零样本泛化能力显著优于基线方法
- 适合研究具身智能、机器人规划与可解释推理的学者
尽管进展迅速,具身智能体在需要保持空间一致性、因果依赖和目标约束的长时程操作任务中仍表现不佳。现有方法将任务推理隐含于高维潜在表示中,难以分离任务结构与感知变化。我们提出地面场景图推理(GSR),一种显式建模世界状态演变为语义基础场景图转换的结构化推理范式。通过分步推理物体状态与空间关系,而非直接从感知映射到动作,GSR能够在物理意义上显式分析动作前提、后果与目标达成。为支持此类推理学习,我们构建了大规模数据集Manip-Cognition-1.6M,联合监督世界理解、动作规划与目标解读。在RLBench、LIBERO、GSR-benchmark及真实机器人任务上的广泛评估表明,GSR显著提升了零样本泛化能力与长时程任务完成率,超越基于提示的基线方法。结果凸显显式世界状态表征作为可扩展具身推理的关键归纳偏置。
原文摘要 · Abstract (English)
Despite rapid progress, embodied agents still struggle with long-horizon manipulation that requires maintaining spatial consistency, causal dependencies, and goal constraints. A key limitation of existing approaches is that task reasoning is implicitly embedded in high-dimensional latent representations, making it challenging to separate task structure from perceptual variability. We introduce Grounded Scene-graph Reasoning (GSR), a structured reasoning paradigm that explicitly models world-state evolution as transitions over semantically grounded scene graphs. By reasoning step-wise over object states and spatial relations, rather than directly mapping perception to actions, GSR enables explicit reasoning about action preconditions, consequences, and goal satisfaction in a physically grounded space. To support learning such reasoning, we construct Manip-Cognition-1.6M, a large-scale dataset that jointly supervises world understanding, action planning, and goal interpretation. Extensive evaluations across RLBench, LIBERO, GSR-benchmark, and real-world robotic tasks show that GSR significantly improves zero-shot generalization and long-horizon task completion over prompting-based baselines. These results highlight explicit world-state representations as a key inductive bias for scalable embodied reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。