为机器人操作泛化能力建立可复现的评估体系
A Taxonomy for Evaluating Generalist Robot Manipulation Policies
- 提出STAR-Gen分类框架,从视觉、语义、行为三方面定义泛化类型
- 在真实场景中验证发现开源模型在语义泛化上表现不佳
- 提供可复现的评测指南,适合研究泛化能力的开发者
机器人操作的机器学习有望实现对新任务和环境的泛化能力。但如何衡量这些策略向泛化的进展?当前评估泛化仍处于无标准状态,各研究采用不同方式测量,且设置难以复现。本文旨在(1)以细粒度方式系统梳理机器人操作中重要的泛化形式;(2)提供可复现的泛化评估指南。我们首先提出STAR-Gen,一个围绕视觉、语义和行为泛化构建的机器人操作泛化分类体系。随后,通过两个真实世界基准案例进行实例化:一是基于开源模型与Bridge V2数据集,二是基于双臂ALOHA 2平台,涵盖更灵巧和长时程任务。案例揭示关键洞见:尽管开源视觉-语言-动作模型在互联网规模语料上预训练,但在语义泛化上仍表现薄弱。补充视频及其他材料可在stargen-taxonomy.github.io获取。
原文摘要 · Abstract (English)
Machine learning for robot manipulation promises to unlock generalization to novel tasks and environments. But how should we measure the progress of these policies towards generalization? Evaluating and quantifying generalization is the Wild West of modern robotics, with each work proposing and measuring different types of generalization in their own, often difficult to reproduce settings. In this work, our goal is (1) to outline the forms of generalization we believe are important for robot manipulation in a comprehensive and fine-grained manner, and (2) to provide reproducible guidelines for measuring these notions of generalization. We first propose STAR-Gen, a taxonomy of generalization for robot manipulation structured around visual, semantic, and behavioral generalization. Next, we instantiate STAR-Gen with two case studies on real-world benchmarking: one based on open-source models and the Bridge V2 dataset, and another based on the bimanual ALOHA 2 platform that covers more dexterous and longer horizon tasks. Our case studies reveal many interesting insights: for example, we observe that open-source vision-language-action models often struggle with semantic generalization, despite pre-training on internet-scale language datasets. We provide videos and other supplementary material at stargen-taxonomy.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。