arXiv:2606.10382cs.RO2026-06被引 1

首个面向UMI机械臂操作的可复现真实世界评测基准

UMI-Bench 1.0: An Open and Reproducible Real-World Benchmark for Tabletop Robotic Manipulation with UMI Data

论文配图:UMI-Bench 1.0: An Open and Reproducible Real-World Benchmark for Tabletop Robotic Manipulation with UMI Data
图 1 · 摘自论文原文
  • 构建统一协议,整合数据采集到任务分析全流程
  • 支持真实场景下UMI策略泛化能力标准化评估
  • 适合研究机器人通用操作与真实部署的团队使用

真实机器人评估对于判断学习的操控策略能否在非受控环境中可靠运行至关重要。这一需求对依赖腕部视角观测、动作表示、数据采集与物理部署耦合的通用操控接口(UMI)类策略尤为突出。现有真实世界基准虽有进展,但未围绕UMI数据到部署的设定设计。本文提出UMI-Bench 1.0,一个以本地优先为原则的真实机器人基准,用于标准化评估UMI风格操控策略。据我们所知,这是首个专注于真实世界评估UMI基操控模型的基准。UMI-Bench在统一协议下对齐数据采集、场景重置、策略执行、结果记录和任务因子分析。通过使整个评估过程可复现、可审计,该基准为衡量UMI训练策略在真实物理操控中的泛化能力提供了实用测试平台。

原文摘要 · Abstract (English)

Real-robot evaluation is essential for understanding whether learned manipulation policies can operate reliably outside curated demonstrations. This need is particularly pressing for Universal Manipulation Interface (UMI)-style policies, whose performance depends on the coupling between wrist-view observations, action representation, data collection, and physical deployment. Existing real-world benchmarks have made important progress, but they are not designed around this UMI data-to-deployment setting. We present UMI-Bench 1.0, a local-first real-robot benchmark for standardized evaluation of UMI-style manipulation policies. To the best of our knowledge, this is the first benchmark dedicated to real-world evaluation of UMI-based manipulation models. UMI-Bench aligns data collection, scene reset, policy execution, result logging, and task-factor analysis within a unified protocol. By making the full evaluation process reproducible and auditable, UMI-Bench provides a practical testbed for measuring how UMI-trained policies generalize to real physical manipulation.

机器人操控真实世界评估可复现性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。