arXiv:2604.09860cs.ROcs.AI2026-04被引 14

构建高保真仿真基准,评估机器人通用策略的真实泛化能力

RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies

论文配图:RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
图 1 · 摘自论文原文
  • 通过人工与大模型生成任务场景,实现跨任务、跨环境的无偏测试
  • 120个任务覆盖视觉/流程/关系三类能力,分三级难度,揭示真实性能差距
  • 量化策略对扰动的敏感度,适合评估前沿机器人模型的鲁棒性

通用机器人研究虽已取得显著进展,但仿真基准仍面临性能迅速饱和和缺乏真正泛化测试的问题。现有基准常存在训练与评估域重叠,导致成功率虚高,难以反映真实鲁棒性。本文提出RoboLab,一个面向任务通用型策略的仿真基准框架,旨在回答两个问题:(1)能否通过仿真行为推断真实策略性能;(2)哪些因素最影响策略表现。首先,该框架支持在高保真仿真环境中以机器人与策略无关方式生成场景与任务。我们构建了RoboLab-120基准,包含120项任务,按视觉、流程、关系三类能力轴划分,覆盖三个难度等级。其次,我们系统分析真实世界策略,量化其性能及对可控扰动的敏感性,揭示当前顶尖模型存在显著性能差距。通过提供细粒度指标与可扩展工具集,RoboLab为评估任务通用型机器人策略的真实泛化能力提供了可靠框架。

原文摘要 · Abstract (English)

The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck due to rapid performance saturation and a lack of true generalization testing. Existing benchmarks often exhibit significant domain overlap between training and evaluation, trivializing success rates and obscuring insights into robustness. We introduce RoboLab, a simulation benchmarking framework designed to address these challenges. Concretely, our framework is designed to answer two questions: (1) to what extent can we understand the performance of a real-world policy by analyzing its behavior in simulation, and (2) which factor most strongly affect policy behavior. First, RoboLab enables human-authored and LLM-enabled generation of scenes and tasks in a robot- and policy-agnostic manner within a high-fidelity simulation environment. We introduce an accompanying RoboLab-120 benchmark, consisting of 120 tasks categorized into three competency axes: visual, procedural, relational, across three difficulty levels. Second, we introduce a systematic analysis of real-world policies that quantify both their performance and the sensitivity of their behavior to controlled perturbations, exposing significant performance gap in current state-of-the-art models. By providing granular metrics and a scalable toolset, RoboLab offers a scalable framework for evaluating the true generalization capabilities of task-generalist robotic policies. Project website: https://research.nvidia.com/labs/srl/projects/robolab/.

机器人仿真基准泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。