arXiv:2607.14439cs.LGcs.RO2026-07

用主动实验设计高效评估机器人泛化能力,少试20%-40%次数就能发现薄弱环节。

Active Real-World Factor-Based Evaluation for Generalist Robot Policies

论文配图:Active Real-World Factor-Based Evaluation for Generalist Robot Policies
图 1 · 摘自论文原文
  • 将评估任务建模为序列实验设计,自适应选择最能获取信息的测试配置。
  • 在3个任务上仅需2331次真实测试,就比随机测试节省20%-40%样本量。
  • 适合需要高效验证复杂场景下机器人策略稳定性的研究者与工程师。

在大规模多样数据集上训练的通用机器人操作策略在多种任务中展现出显著潜力。然而,对这些策略进行严谨评估仍是根本挑战。真实世界表现依赖于物体姿态、相机视角等大量组合因素,全量遍历评估不可行;且硬件测试耗时耗力,当前常用窄范围测试集易遗漏关键失败模式,误判部署准备度。本文提出一种主动评估框架,将策略评估视为序列实验设计问题。通过在结构化任务因子空间上构建概率代理模型,自适应选择能最大化信息增益的评估配置,实现对未见条件下的策略行为的高样本效率刻画,并系统识别出易失效区域。我们在3个任务、3类因子变化下进行了2331次真实世界评估,结果表明该方法通常可比常规随机测试减少20%-40%的试验次数。

原文摘要 · Abstract (English)

Generalist robot manipulation policies trained on large, diverse datasets have shown remarkable promise across a wide range of tasks. However, rigorously evaluating these policies remains a fundamental challenge. Real-world performance depends on a large combinatorial space of task factors including object poses and camera viewpoints, making full, exhaustive evaluation intractable. Additionally, real hardware evaluation is slow and resource-intensive, so current practice is to use narrow test suites that can miss critical failure modes and misrepresent true deployment readiness. We propose an active evaluation framework that addresses this challenge by treating policy evaluation as a sequential experimental design problem. Our approach fits a probabilistic surrogate model over a structured space of task factors and adaptively selects evaluation configurations to maximize information gain over the policy's performance distribution, allowing for sample-efficient characterization of policy behavior across unseen conditions and a systematic identification of failure-prone regions. We conduct 2331 real-world evaluations across 3 tasks with 3 factor variations and find that our approach typically saves the evaluator at least 20-40% of trials compared to typical random testing.

机器人评估主动学习策略验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。