从信息论角度揭示模仿学习泛化能力的瓶颈与优化方向
Generalization Capability for Imitation Learning
- 用条件信息瓶颈和参数-数据互信息约束泛化误差上界
- 高输入输出熵可降低泛化差距并加速跳出尖锐极小值
- 适合关注机器人技能泛化、模型训练策略设计的研究者
模仿学习通过专家示范赋予机器人多样化技能,但有限数据训练的策略常难以超越训练分布。本文从信息论与数据分布特性出发,提出统一视角:泛化差距可被中间表征的条件信息瓶颈和模型参数与训练数据间的互信息所上界约束。该理论指导训练策略设计,尤其在决定是否冻结、微调或从头训练大模型(如视觉-语言模型)时具有参考价值。此外,输入到输出的高条件熵使似然景观更平坦,从而缩小泛化差距上界,并缩短随机梯度下降(SGD)从尖锐局部极小值逃逸的时间,提升在固定优化预算下达到全局最优的可能性。这些发现解释了模仿学习泛化受限的原因,强调不仅需扩大输入数据多样性,还需增强相同输入下的输出标签变异性。
原文摘要 · Abstract (English)
Imitation learning holds the promise of equipping robots with versatile skills by learning from expert demonstrations. However, policies trained on finite datasets often struggle to generalize beyond the training distribution. In this work, we present a unified perspective on the generalization capability of imitation learning, grounded in both information theorey and data distribution property. We first show that the generalization gap can be upper bounded by (i) the conditional information bottleneck on intermediate representations and (ii) the mutual information between the model parameters and the training dataset. This characterization provides theoretical guidance for designing effective training strategies in imitation learning, particularly in determining whether to freeze, fine-tune, or train large pretrained encoders (e.g., vision-language models or vision foundation models) from scratch to achieve better generalization. Furthermore, we demonstrate that high conditional entropy from input to output induces a flatter likelihood landscape, thereby reducing the upper bound on the generalization gap. In addition, it shortens the stochastic gradient descent (SGD) escape time from sharp local minima, which may increase the likelihood of reaching global optima under fixed optimization budgets. These insights explain why imitation learning often exhibits limited generalization and underscore the importance of not only scaling the diversity of input data but also enriching the variability of output labels conditioned on the same input.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。