arXiv:2506.04399cs.LGcs.AI2025-06被引 7

无需奖励信号,用少量交互快速适应新任务的元强化学习方法。

Unsupervised Meta-Testing with Conditional Neural Processes for Hybrid Meta-Reinforcement Learning

  • 用条件神经过程离线学习任务推断,复用已有数据提升效率。
  • 单次测试轨迹即可推断动态模型,减少90%以上在线交互次数。
  • 适合无奖励信号、样本稀缺的现实强化学习场景。

我们提出一种无监督元测试方法UMCNP,结合参数化策略梯度与任务推断的少样本元强化学习。该方法在元测试阶段无奖励信号时仍能高效适应,无需额外元训练样本。通过条件神经过程(CNPs)高效复用元训练中已收集的样本,实现离线任务推断。在测试时仅需单次任务轨迹,即可推断转移动态的潜在表示,并基于学习的动态模型生成自适应滚动数据。在2D-Point Agent和连续控制基准任务(如含未知角度传感器偏差的CartPole、动态参数随机化的Walker代理)上,相比基线方法,元测试阶段样本使用量显著减少。

原文摘要 · Abstract (English)

We introduce Unsupervised Meta-Testing with Conditional Neural Processes (UMCNP), a novel hybrid few-shot meta-reinforcement learning (meta-RL) method that uniquely combines, yet distinctly separates, parameterized policy gradient-based (PPG) and task inference-based few-shot meta-RL. Tailored for settings where the reward signal is missing during meta-testing, our method increases sample efficiency without requiring additional samples in meta-training. UMCNP leverages the efficiency and scalability of Conditional Neural Processes (CNPs) to reduce the number of online interactions required in meta-testing. During meta-training, samples previously collected through PPG meta-RL are efficiently reused for learning task inference in an offline manner. UMCNP infers the latent representation of the transition dynamics model from a single test task rollout with unknown parameters. This approach allows us to generate rollouts for self-adaptation by interacting with the learned dynamics model. We demonstrate our method can adapt to an unseen test task using significantly fewer samples during meta-testing than the baselines in 2D-Point Agent and continuous control meta-RL benchmarks, namely, cartpole with unknown angle sensor bias, walker agent with randomized dynamics parameters.

元强化学习少样本学习无监督测试动态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。