arXiv:2608.13415cs.ROcs.AI2026-08

在有限练习时间下,让机器人高效学习实用技能。

Deliberate Practice: Learning Robot Skills under a Budget

论文配图:Deliberate Practice: Learning Robot Skills under a Budget
图 1 · 摘自论文原文
  • 设计算法动态分配练习时间,优先学收益高且可掌握的技能。
  • 实验证明能在有限时间内提升长程任务规划能力。
  • 适合资源受限的机器人自主学习场景。

我们研究在有限实践预算下,自主学习机器人序列任务技能的问题。提出一种主动技能学习算法——刻意练习(Deliberate Practice, DP),能计算出理论上最优的预算分配:选择那些在预算内可掌握、且能带来最大累积奖励的技能进行练习。DP 同时估算掌握技能所需时间以及由此解锁的任务计划的累积奖励。由于需在大预算范围内考虑组合爆炸式技能方案,计算最优分配极具挑战性。我们的核心贡献是构建了一个双线性规划模型,可借助现成求解器精确求解。在长程操作任务的仿真与真实世界实验中,结果表明该方法能让机器人在有限练习时间内,最优地获取有效策略并提升长程规划性能。

原文摘要 · Abstract (English)

We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tasks. We propose an active skill learning algorithm, \emph{Deliberate Practice (DP)}, that computes a provably \emph{budget-optimal} allocation---practicing skills that maximize expected cumulative reward while being learnable within the budget. DP estimates both the time needed to master skills and the cumulative reward of the task plans that the skills unlock. Computing a budget-optimal allocation is challenging as it requires reasoning about combinatorially many skill plans over a large practice budget. Our key contribution is a bilinear program that can compute this exactly using off-the-shelf solvers. Through simulated and real-world experiments on long-horizon manipulation tasks, we show that our approach allows robots to optimally use limited practice time to acquire useful policies and improve long-horizon planning.

机器人学习预算优化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。