arXiv:2501.15270cs.LGcs.AI2025-01

用语言引导强化学习,提升零样本系统泛化与采样效率

Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning

  • 通过神经生成系统构建模块化稀疏表示,增强语言与决策关联
  • 在BabyAI上实现比以往模型更优的系统泛化和采样效率
  • 记忆模块提供高层信息聚合与注意力反馈,提升训练稳定性

样本效率和系统泛化是强化学习长期面临的挑战。以往研究发现,将自然语言与其他观测模态结合可因语言的组合性与开放性提升泛化能力与样本效率。但要将语言特性有效传递至决策过程,需建立合适的语言接地机制。本文提出基于神经生成系统(NPS)的架构级归纳偏置,强调模块化与稀疏性,并赋予记忆组件核心地位。记忆作为高层信息聚合器,向策略/价值头提供综合信息,同时通过注意力反馈引导NPS的选通注意。在BabyAI环境中的实验表明,该模型在系统泛化和样本效率上显著优于先前方法。通过广泛的消融实验,明确了各项技术对泛化、采样效率及训练稳定性的具体贡献。

原文摘要 · Abstract (English)

Sample efficiency and systematic generalization are two long-standing challenges in reinforcement learning. Previous studies have shown that involving natural language along with other observation modalities can improve generalization and sample efficiency due to its compositional and open-ended nature. However, to transfer these properties of language to the decision-making process, it is necessary to establish a proper language grounding mechanism. One approach to this problem is applying inductive biases to extract fine-grained and informative representations from the observations, which makes them more connectable to the language units. We provide architecture-level inductive biases for modularity and sparsity mainly based on Neural Production Systems (NPS). Alongside NPS, we assign a central role to memory in our architecture. It can be seen as a high-level information aggregator which feeds policy/value heads with comprehensive information and simultaneously guides selective attention in NPS through attentional feedback. Our results in the BabyAI environment suggest that the proposed model's systematic generalization and sample efficiency are improved significantly compared to previous models. An extensive ablation study on variants of the proposed method is conducted, and the effectiveness of each employed technique on generalization, sample efficiency, and training stability is specified.

强化学习语言引导系统泛化神经生成系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。