arXiv:2508.13998cs.ROcs.AI2025-08被引 45

用‘指’作为通用中间表示,让机器人更懂视觉指令并精准执行。

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation

  • 以‘指’为统一中间表示,连接视觉理解与动作执行。
  • 零样本在真实机械臂任务中达87.5%成功率,比基线高62%。
  • 适合需要强泛化能力的具身智能与机器人控制研究者。

具身人工智能的泛化能力受限于“看见到执行”的鸿沟,源于数据稀缺与具身异构性。为此,我们首次将“指”定义为一种无具身依赖的统一中间表示,提出四项核心具身指代表达能力,弥合高层视觉语言理解与底层动作原语之间的差距。我们构建了30亿参数的具身推理视觉语言模型Embodied-R1,并基于多源具身与通用视觉推理数据集创建大规模数据集Embodied-Points-200K,支持关键指代表达能力。采用两阶段强化微调(RFT)训练范式,设计多任务奖励机制。Embodied-R1在11个具身空间与指代基准上达到顶尖性能,尤其在无任务微调条件下,于SIMPLEREnv中实现56.2%成功率,在8项真实世界XArm任务中平均达87.5%,相较强基线提升62%。模型对多样视觉干扰也表现出高度鲁棒性。结果表明,以指为核心表征结合RFT训练范式,是有效闭合机器人感知-动作差距的通用路径。

原文摘要 · Abstract (English)

Generalization in embodied AI is hindered by the "seeing-to-doing gap," which stems from data scarcity and embodiment heterogeneity. To address this, we pioneer "pointing" as a unified, embodiment-agnostic intermediate representation, defining four core embodied pointing abilities that bridge high-level vision-language comprehension with low-level action primitives. We introduce Embodied-R1, a 3B Vision-Language Model (VLM) specifically designed for embodied reasoning and pointing. We use a wide range of embodied and general visual reasoning datasets as sources to construct a large-scale dataset, Embodied-Points-200K, which supports key embodied pointing capabilities. We then train Embodied-R1 using a two-stage Reinforced Fine-tuning (RFT) curriculum with a specialized multi-task reward design. Embodied-R1 achieves state-of-the-art performance on 11 embodied spatial and pointing benchmarks. Critically, it demonstrates robust zero-shot generalization by achieving a 56.2% success rate in the SIMPLEREnv and 87.5% across 8 real-world XArm tasks without any task-specific fine-tuning, representing a 62% improvement over strong baselines. Furthermore, the model exhibits high robustness against diverse visual disturbances. Our work shows that a pointing-centric representation, combined with an RFT training paradigm, offers an effective and generalizable pathway to closing the perception-action gap in robotics.

具身智能视觉语言模型机器人操控强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。