用语言批评指导学习,从差示范中学会正确动作
Language-Critique Imitation Learning from Suboptimal Demonstrations

- 用自然语言描述进展、错误和修正,构建结构化监督信号
- 在多个连续控制任务中超越主流模仿学习与离线强化学习方法
- 适合需要从真实差数据中学习的智能体训练场景
以往从次优示范中进行模仿学习的方法通常依赖置信度、判别器分数或重要性权重等压缩型标量信号,这类信号难以表达任务进展、失败原因或纠正动作等中间推理。本文提出一种基于自然语言的批判式模仿学习框架,利用语言作为结构化监督信号,避免反馈信息被简化为标量。方法首先从示范中构建包含当前进展、次优行为识别和细粒度修正指导的语言标签;随后引入语言批判损失,直接使用这些结构化信号训练策略,无需降维为标量,并分别应用于行为克隆与扩散策略,得到LC-BC和LC-DP。我们进一步提供理论结果表明,在标准假设下,该目标可上界专家性能差距。实验在涵盖导航、操作和游戏的多样化连续控制任务中验证,结果表明我们的方法持续优于强基线,证明语言可作为从次优数据中学习鲁棒策略的强大且结构化的监督形式。
原文摘要 · Abstract (English)
Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals such as confidence estimates, discriminator scores, or importance weights. These scalar signals are inherently limited, as they cannot explicitly express intermediate reasoning about task progress, failure modes, or corrective actions. We propose a language-critique framework for imitation learning from suboptimal demonstrations that instead leverages natural language as a structured supervision signal, avoiding the collapse of expressive feedback into scalars. Our method first constructs language labels from demonstrations that explicitly describe current progress, identify suboptimal behaviors, and provide fine-grained corrective guidance. We then introduce a language-critique loss that directly trains policies using these structured signals without reducing them to scalars, and instantiate it for both behavior cloning and diffusion policies, yielding LC-BC and LC-DP. We further provide a theoretical result showing that the proposed objective upper-bounds the expert performance gap under standard assumptions. Empirically, we evaluate on diverse continuous control tasks spanning navigation, manipulation, and gameplay, where our methods consistently outperform strong imitation learning and offline reinforcement learning baselines. These results demonstrate that language can serve as a powerful and structured form of supervision for learning robust policies from suboptimal data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。