arXiv:2510.23448cs.LGstat.ML2025-10

从信息论角度分析元强化学习在分布外场景下的泛化能力。

An Information-Theoretic Analysis of OOD Generalization in Meta-Reinforcement Learning

  • 用信息论构建元监督学习的泛化上界,涵盖分布偏移与宽到窄训练场景。
  • 针对马尔可夫决策过程结构,给出细粒度的元强化学习泛化边界。
  • 分析梯度法元强化学习算法的泛化性能,揭示其内在机制。

本文从信息论视角研究元强化学习中的分布外(OOD)泛化问题。首先,在标准分布不匹配和广义到狭义训练设置两种分布偏移场景下,建立元监督学习的泛化上界。在此基础上,形式化元强化学习的泛化问题,并利用马尔可夫决策过程的结构特性,推导出更精细的泛化边界。最后,分析基于梯度的元强化学习算法的泛化表现,为理解其在未见任务上的鲁棒性提供理论依据。

原文摘要 · Abstract (English)

In this work, we study out-of-distribution (OOD) generalization in meta-reinforcement learning from an information-theoretic perspective. We begin by establishing OOD generalization bounds for meta-supervised learning under two distinct distribution shift scenarios: standard distribution mismatch and a broad-to-narrow training setting. Building on this foundation, we formalize the generalization problem in meta-reinforcement learning and establish fine-grained generalization bounds that exploit the structure of Markov Decision Processes. Lastly, we analyze the generalization performance of a gradient-based meta-reinforcement learning algorithm.

元强化学习泛化分析信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。