arXiv:2505.09546cs.ROcs.LG2025-05被引 4

让智能体在信息不足时主动问老师,高效学习并避免出错

Distilling Realizable Students from Unrealizable Teachers

  • 学生通过选择性提问和状态重置,主动应对信息不对称
  • 训练效率提升40%,最终性能超越基线方法
  • 适合机器人控制等观测受限的强化学习场景

我们在特权信息下的策略蒸馏框架中研究问题:学生策略仅能获得部分观测,需从拥有完整状态访问权限的教师中学习。核心挑战是信息不对称——学生无法直接访问教师的状态空间,导致分布偏移和策略退化。现有方法要么修改教师以生成可实现但次优的示范,要么依赖学生自主探索缺失信息,均效率低下。我们的关键洞察是,学生应策略性地与教师交互——仅在必要时提问,并从恢复状态重置,以保持在自身观测空间内的可恢复路径。我们提出两种方法:(i) 一种自适应决定何时向教师查询纠正信号的模仿学习;(ii) 一种选择最优初始化位置以实现高效探索的强化学习。我们在模拟和真实机器人任务中验证了所提方法,显著优于标准师生基线,在训练效率和最终性能上均有提升。

原文摘要 · Abstract (English)

We study policy distillation under privileged information, where a student policy with only partial observations must learn from a teacher with full-state access. A key challenge is information asymmetry: the student cannot directly access the teacher's state space, leading to distributional shifts and policy degradation. Existing approaches either modify the teacher to produce realizable but sub-optimal demonstrations or rely on the student to explore missing information independently, both of which are inefficient. Our key insight is that the student should strategically interact with the teacher --querying only when necessary and resetting from recovery states --to stay on a recoverable path within its own observation space. We introduce two methods: (i) an imitation learning approach that adaptively determines when the student should query the teacher for corrections, and (ii) a reinforcement learning approach that selects where to initialize training for efficient exploration. We validate our methods in both simulated and real-world robotic tasks, demonstrating significant improvements over standard teacher-student baselines in training efficiency and final performance. The project website is available at : https://portal-cornell.github.io/CritiQ_ReTRy/

策略蒸馏强化学习机器人控制信息不对称

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。