arXiv:2509.25466cs.LGcs.AI2025-09

用少样本数据训练通用机器人策略,自动聚焦弱项任务提升成功率。

Data-Efficient Multitask DAgger

  • 通过性能感知调度,动态分配示范数据给表现差的任务。
  • 在MetaWorld和IsaacLab上用更少演示达到高成功率,零样本迁移真实机器人效果更好。
  • 适合资源有限、需跨任务泛化的机器人学习研究者。

通用机器人策略通常需要大量专家数据或仿真训练。本文提出一种数据高效的多任务DAgger框架,从多个任务专用专家策略中提炼单一多任务策略。该方法通过主动聚焦表现不佳的任务,显著提升整体任务成功率。核心是基于卡尔曼滤波器的性能感知调度策略,可稳健评估各任务学习收益并决定示范数据分配。我们在MetaWorld及IsaacLab的多样化抽屉开启任务上验证了该方法。结果表明,所学策略在所有任务上均表现优异,且所需专家演示大幅减少;在仿真中训练的视觉策略无需真实数据,零样本迁移到真实机器人时表现优于朴素DAgger和行为克隆。

原文摘要 · Abstract (English)

Generalist robot policies that can perform many tasks typically require extensive expert data or simulations for training. In this work, we propose a novel Data-Efficient multitask DAgger framework that distills a single multitask policy from multiple task-specific expert policies. Our approach significantly increases the overall task success rate by actively focusing on tasks where the multitask policy underperforms. The core of our method is a performance-aware scheduling strategy that tracks how much each task's learning process benefits from the amount of data, using a Kalman filter-based estimator to robustly decide how to allocate additional demonstrations across tasks. We validate our approach on MetaWorld, as well as a suite of diverse drawer-opening tasks in IsaacLab. The resulting policy attains high performance across all tasks while using substantially fewer expert demonstrations, and the visual policy learned with our method in simulation shows better performance than naive DAgger and Behavior Cloning when transferring zero-shot to a real robot without using real data.

机器人学习多任务数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。