arXiv:2602.06550cs.LGcs.AI2026-02

用共享超网络实现上下文突变时的零样本泛化,提升强化学习鲁棒性。

Dynamics-Aligned Shared Hypernetworks for Contextual RL under Discontinuous Shifts

  • 单个超网络通过预测环境动态生成适配器权重,统一调节策略与价值函数。
  • 在新任务上零样本表现超越领域随机化58.1%,优于基线11.5%。
  • 适用于上下文突变、隐含状态的复杂强化学习场景,如机械臂控制。

上下文强化学习中的零样本泛化仍是核心挑战,尤其当上下文为隐含状态且需从数据中推断时。当隐含上下文发生突变导致动作对环境的影响方式改变时,传统方法易失效。本文提出DMA*-SH框架:一个仅通过环境动态预测训练的超网络,生成一组共享的适配器权重,用于动态模型、策略和动作价值函数。该共享调制引入与上下文-动态突变匹配的归纳偏置;输入/输出归一化与随机输入掩码增强上下文推理,促进方向集中表征。理论方面,提供了超网络调制的表达能力分离结果,以及基于策略梯度方差的分解分析,证明模式内压缩可提升非重叠上下文下的学习效率。我们构建了执行器反演基准(AIB),包含执行器反演、排列及弱非重叠连续动态等任务。在AIB的未见任务上,DMA*-SH实现零样本泛化,平均优于领域随机化58.1%,优于标准上下文感知基线11.5%。

原文摘要 · Abstract (English)

Zero-shot generalization in contextual reinforcement learning remains a core challenge, particularly when the context is latent and must be inferred from data. A canonical failure mode arises when latent context discontinuously changes how actions affect the environment, requiring incompatible control responses across contexts. We propose DMA*-SH, a framework where a single hypernetwork, trained solely via dynamics prediction, generates a small set of adapter weights shared across the dynamics model, policy, and action-value function. This shared modulation imparts an inductive bias matched to discontinuous context-to-dynamics shifts, while input/output normalization and random input masking stabilize context inference, promoting directionally concentrated representations. We provide theoretical support via expressivity separation results for hypernetwork modulation, and a variance decomposition with policy-gradient variance bounds that formalize how within-mode compression improves learning under non-overlapping contexts. For evaluation, we introduce the Actuator Inversion Benchmark (AIB), a suite of environments designed to isolate challenging context-to-dynamics interactions, including actuator inversion, actuator permutations, and weakly non-overlapping continuous dynamics. On AIB's held-out tasks, DMA*-SH achieves zero-shot generalization, outperforming domain randomization by 58.1% and surpassing a standard context-aware baseline by 11.5% on average.

强化学习零样本上下文泛化超网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。