arXiv:2608.18433cs.ROcs.LG2026-08

揭示机器人模型在不同身体上迁移时的适配鸿沟,指导高效部署。

The Embodiment Gap in Robot Foundation Models

  • 提出'具身鸿沟'概念,分析可复用与需重做的部分。
  • 构建二维地图,定位方法在共享结构与适配阶段的位置。
  • 建议新评估框架,超越成功率看实际适配工作量。

机器人基础模型(RFM),包括视觉-语言-动作(VLA)策略,常被以扩展视角讨论:更多数据、更大模型、更广基准应提升泛化能力。然而,在机器人领域,模型虽能泛化,仍需大量工作才能在特定机器人身体上运行。这些工作量因方法和目标机器人而异,影响实际部署。我们称模型、表征或数据在目标机器人执行前的可用性差距为具身鸿沟。本综述分析了跨机器人身体可复用的内容及必须在新机器人上实现的部分。通过二维坐标图展示共享结构类型与适配阶段。进一步从三个重叠方向审视近期工作:共享语义与感知、共享机器人数据与接口、跨身体对应关系学习。还提出一个报告框架,用于揭示适应工作量,仅靠成功率无法体现。该框架明确比较跨身体学习时应检查的工作,并指出新机器人仍需完成的任务及未来研究问题。

原文摘要 · Abstract (English)

Robot foundation models (RFMs), including vision-language-action (VLA) policies, are often discussed through a scaling view: more data, larger models, and broader benchmarks should improve generalization. In robotics, however, a model can generalize while work still remains before it can run on a robot with a particular body. The work required differs across methods and target robots, and those differences affect practical deployment. We call the gap between reusable models, representations, or data and their use in execution on the target robot the embodiment gap. This survey examines what can be reused across robot embodiments and what must still be implemented on a new robot. We place existing methods on a two-axis map that shows the type of shared structure and the stage at which adaptation is needed for execution on the target robot. We then examine recent work through three overlapping research directions: sharing semantics and perception, sharing robot data and interfaces, and learning correspondence across embodiments. We also propose a reporting framework for adaptation work that success rate alone does not reveal. The framework identifies the work that should be checked when comparing cross-embodiment learning and highlights work that remains on a new robot and questions for future study.

机器人具身智能迁移学习基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。