arXiv:2502.09829cs.ROcs.AI2025-02被引 7

用主动实验选择减少机器人多任务评估成本

Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection

  • 基于任务间相似性与自然语言先验,动态选最有信息量的实验
  • 在真实机器人和仿真数据上降低评估成本,保持结果可靠
  • 适合需要高效评估大量策略与任务的机器人研发团队

评估学习到的机器人控制策略在物理任务中的能力需耗费大量人力。随着策略与任务数量增长,全面测试变得不切实际——每次试验需手动重置环境,任务切换常需重新布置物体甚至更换机器人。盲目随机选取部分组合测试成本高且结果不可靠。本文将机器人评估建模为一种主动测试问题:通过序列化执行实验,逐步构建策略-任务性能分布模型。利用任务间的相似性及自然语言作为先验知识,揭示策略行为的潜在关联。据此设计成本感知的期望信息增益启发式方法,高效筛选具有信息量的试验。框架支持连续与离散性能输出。在真实机器人与仿真数据上的实验表明,该方法显著降低跨多任务评估机器人策略的计算成本。

原文摘要 · Abstract (English)

Evaluating learned robot control policies to determine their physical task-level capabilities costs experimenter time and effort. The growing number of policies and tasks exacerbates this issue. It is impractical to test every policy on every task multiple times; each trial requires a manual environment reset, and each task change involves re-arranging objects or even changing robots. Naively selecting a random subset of tasks and policies to evaluate is a high-cost solution with unreliable, incomplete results. In this work, we formulate robot evaluation as an active testing problem. We propose to model the distribution of robot performance across all tasks and policies as we sequentially execute experiments. Tasks often share similarities that can reveal potential relationships in policy behavior, and we show that natural language is a useful prior in modeling these relationships between tasks. We then leverage this formulation to reduce the experimenter effort by using a cost-aware expected information gain heuristic to efficiently select informative trials. Our framework accommodates both continuous and discrete performance outcomes. We conduct experiments on existing evaluation data from real robots and simulations. By prioritizing informative trials, our framework reduces the cost of calculating evaluation metrics for robot policies across many tasks.

机器人评估主动学习多任务实验优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。