用自由概率理论生成层级结构随机投影,提升强化学习泛化能力
Free Random Projection for In-Context Reinforcement Learning
- 基于自由概率构造随机正交矩阵,自然产生层次结构
- 在多环境基准上优于标准随机投影,泛化性能显著提升
- 无需修改架构即可融入现有框架,适合复杂状态空间任务
层次化归纳偏置被认为能促进强化学习中的可泛化策略,已有研究通过显式的双曲潜空间和网络结构验证了这一点。为此,我们提出一种更灵活的方法:让这些偏置从算法中自然涌现。本文引入自由随机投影(Free Random Projection),其基于自由概率论构建随机正交矩阵,使层次结构自然生成。该方法可无缝集成至现有上下文强化学习框架中,通过在输入空间编码层次组织,无需显式架构修改。多环境基准上的实验表明,自由随机投影持续优于标准随机投影,提升了泛化性能。在线性可解马尔可夫决策过程分析及核随机矩阵谱的考察中,揭示了其性能增强的理论基础,凸显其在层次化状态空间中的高效适应能力。
原文摘要 · Abstract (English)
Hierarchical inductive biases are hypothesized to promote generalizable policies in reinforcement learning, as demonstrated by explicit hyperbolic latent representations and architectures. Therefore, a more flexible approach is to have these biases emerge naturally from the algorithm. We introduce Free Random Projection, an input mapping grounded in free probability theory that constructs random orthogonal matrices where hierarchical structure arises inherently. The free random projection integrates seamlessly into existing in-context reinforcement learning frameworks by encoding hierarchical organization within the input space without requiring explicit architectural modifications. Empirical results on multi-environment benchmarks show that free random projection consistently outperforms the standard random projection, leading to improvements in generalization. Furthermore, analyses within linearly solvable Markov decision processes and investigations of the spectrum of kernel random matrices reveal the theoretical underpinnings of free random projection's enhanced performance, highlighting its capacity for effective adaptation in hierarchically structured state spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。