arXiv:2607.24744cs.ROcs.CV2026-07综述被引 1

构建机器人操作数据金字塔,揭示数据来源与模型能力的关系

Data Pyramid for Embodied Manipulation: A Survey

论文配图:Data Pyramid for Embodied Manipulation: A Survey
图 1 · 摘自论文原文
  • 提出五层数据金字塔,整合真实机器人到通用视觉语言数据
  • 发现数据物理保真度与可扩展性存在权衡,影响模型性能
  • 适合研究机器人学习、具身智能系统设计的学者参考

多模态基础模型通过互联网数据实现感知与语言理解,但具身智能体缺乏此类捷径,因其需耦合观测、物理状态与动作的数据。本文将具身数据生态组织为五层金字塔:真实机器人数据、UMI风格数据、第一人称与第三人称数据、仿真数据以及通用视觉-语言数据。金字塔围绕可扩展性与机器人对齐之间的张力构建,并从数据质量、多样性、可复用性与物理保真度五个维度刻画每类数据。我们分析近期具身基础模型的训练数据配方,考察不同数据源在预训练中的选择、对齐与混合方式。无论对于具身大脑模型、视觉-语言-动作模型还是世界-动作模型,均关联数据构成与感知、推理、规划、动作生成及世界预测等能力。最后提出六个开放挑战:构建大规模触觉数据集、采集失败与恢复数据、开发可扩展数据采集流水线、跨具身动作对齐、利用第一人称数据实现灵巧操作,以及设计机器人学习的合理数据配方。期望本工作为下一代具身系统的设计奠定基础。

原文摘要 · Abstract (English)

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.

具身智能数据金字塔机器人学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。