arXiv:2510.24194cs.ROcs.LG2025-10NeurIPS被引 1

让专家闭眼示范,反而能学得更好、泛化更强。

Blindfolded Experts Generalize Better: Insights from Robotic Manipulation and Videogames

  • 用信息受限的盲视专家示范,激发更有效的探索行为。
  • 在机器人插销任务和Procgen游戏上,泛化性能显著提升。
  • 理论证明:信息越少,泛化误差越小,适合数据稀缺场景。

行为克隆是一种从示范中学习序列决策的简单而有效的方法。近年来,它成为物理世界基础模型的核心,但实现泛化需要大量任务的海量示范。通常,全知的真人专家会展示近乎最优的行为。本文提出将部分任务信息隐藏给示范者,使其处于“闭眼”状态,被迫进行非平凡探索以完成任务。我们发现,克隆这种盲视专家比克隆全知专家更能泛化到未见过的任务。我们在真实机器人插销任务(有限人类示范)和Procgen基准游戏上进行了实验。此外,理论分析表明,泛化误差与√(I/m)成正比,其中I表示示范者可获得的任务信息量,m为示范任务数量。理论与实践均表明,在示范任务较少时,克隆盲视专家的泛化能力更优。

原文摘要 · Abstract (English)

Behavioral cloning is a simple yet effective technique for learning sequential decision-making from demonstrations. Recently, it has gained prominence as the core of foundation models for the physical world, where achieving generalization requires countless demonstrations of a multitude of tasks. Typically, a human expert with full information on the task demonstrates a (nearly) optimal behavior. In this paper, we propose to hide some of the task's information from the demonstrator. This ``blindfolded'' expert is compelled to employ non-trivial exploration to solve the task. We show that cloning the blindfolded expert generalizes better to unseen tasks than its fully-informed counterpart. We conduct experiments of real-world robot peg insertion tasks with (limited) human demonstrations, alongside videogames from the Procgen benchmark. Additionally, we support our findings with theoretical analysis, which confirms that the generalization error scales with $\sqrt{I/m}$, where $I$ measures the amount of task information available to the demonstrator, and $m$ is the number of demonstrated tasks. Both theory and practice indicate that cloning blindfolded experts generalizes better with fewer demonstrated tasks. Project page with videos and code: https://sites.google.com/view/blindfoldedexperts/home

行为克隆泛化能力机器人强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。