arXiv:2510.22039cs.AIq-bio.NC2025-10NeurIPS

用预测编码提升元强化学习,让智能体在部分可观测环境中学会可解释的最优信念表示。

Predictive Coding Enhances Meta-RL To Achieve Interpretable Bayes-Optimal Belief Representation Under Partial Observability

  • 将神经科学中的预测编码引入元强化学习,自监督地学习历史压缩表示。
  • 在多种任务中,新方法生成的信念表示更接近贝叶斯最优,且在复杂任务中表现显著更好。
  • 适合研究可解释性强化学习、部分可观测环境下的智能体设计者。

在部分可观测环境中,学习历史的紧凑表示对规划和泛化至关重要。尽管元强化学习(meta-RL)智能体能接近贝叶斯最优策略,但常无法学到紧凑且可解释的贝叶斯最优信念状态。这种表征低效可能限制其适应性和泛化能力。受神经科学中预测编码(即大脑通过预测感官输入实现贝叶斯推断)及深度强化学习中辅助预测目标的启发,我们探究将自监督预测编码模块集成到元强化学习中是否有助于学习贝叶斯最优表示。通过状态机模拟,我们发现:在多种任务中,加入预测模块的元强化学习始终生成更可解释的表示,更逼近贝叶斯最优信念状态,即使两者均能达到最优策略。在需要主动获取信息的挑战性任务中,仅带预测模块的元强化学习成功学习到最优表示与策略,而传统方法因表征学习不足而失败。最后,我们证明更好的表征学习带来更强泛化能力。结果强烈表明,预测学习是智能体在部分可观测环境中实现有效表征学习的关键指导原则。

原文摘要 · Abstract (English)

Learning a compact representation of history is critical for planning and generalization in partially observable environments. While meta-reinforcement learning (RL) agents can attain near Bayes-optimal policies, they often fail to learn the compact, interpretable Bayes-optimal belief states. This representational inefficiency potentially limits the agent's adaptability and generalization capacity. Inspired by predictive coding in neuroscience--which suggests that the brain predicts sensory inputs as a neural implementation of Bayesian inference--and by auxiliary predictive objectives in deep RL, we investigate whether integrating self-supervised predictive coding modules into meta-RL can facilitate learning of Bayes-optimal representations. Through state machine simulation, we show that meta-RL with predictive modules consistently generates more interpretable representations that better approximate Bayes-optimal belief states compared to conventional meta-RL across a wide variety of tasks, even when both achieve optimal policies. In challenging tasks requiring active information seeking, only meta-RL with predictive modules successfully learns optimal representations and policies, whereas conventional meta-RL struggles with inadequate representation learning. Finally, we demonstrate that better representation learning leads to improved generalization. Our results strongly suggest the role of predictive learning as a guiding principle for effective representation learning in agents navigating partial observability.

元强化学习预测编码信念表示部分可观测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。