arXiv:2606.00880cs.LGcs.AI2026-06

任务多样性提升迁移能力但阻碍持续学习,提出新基准Banyan验证其边界。

Task diversity produces systematic transfer but inhibits continual reinforcement learning

论文配图:Task diversity produces systematic transfer but inhibits continual reinforcement learning
图 1 · 摘自论文原文
  • 构建三轴可控任务多样性环境Banyan,可分离测试布局、物体与目标依赖的影响。
  • 多样性提升初始迁移性能,但多轮分布变化后长程任务停滞且旧任务遗忘。
  • 为研究持续强化学习中任务多样性的作用提供可复现的基准测试平台。

持续强化学习旨在训练出既能提升当前任务表现,又能适应任务分布变化的智能体。在多个多样化任务上训练可引发零样本泛化,但以往研究多在训练后冻结权重评估泛化效果。任务多样性是否真正增强智能体跨分布转移的持续学习能力仍不明确。本文提出Banyan——一个基于GPU加速的持续强化学习基准环境,其任务多样性由三个独立可控维度构成:需导航的地图布局、需交互的物体、以及子目标依赖的层次结构。在单次分布变化下,沿任一维度增加多样性均能使智能体在新任务上的起始性能接近前一任务的表现,即使最优策略结构发生改变。然而,随着分布变化次数增多,这种局部迁移无法维持持续学习:长时序任务性能停滞,先前任务分布被后续训练遗忘。Banyan为研究可控任务多样性在何时产生可迁移学习、迁移是否持久、以及何处未能达到真正持续学习提供了系统性评测框架。

原文摘要 · Abstract (English)

Continual reinforcement learning aims to produce agents that learn not only to improve at their current tasks but also to adapt as task distributions change. Training an agent on many diverse tasks can induce zero-shot generalization, but previous work generally evaluates this generalization after training -- with frozen weights. Whether task diversity also improves an agent's ability to continue learning across distribution shifts remains unclear. We introduce Banyan, a GPU-accelerated continual RL domain in which task diversity factors into three independently controllable axes: the map layouts an agent must navigate, the objects it must interact with, and the hierarchical structures of sub-goal dependencies. Across individual distribution shifts, increasing diversity along each axis causes agents to begin training on the new tasks near the performance attained on the previous one, even when the shift changes the structure of the optimal policy. However, as the number of shifts increases, this local transfer does not by itself yield sustained continual learning: longer-horizon tasks plateau, and earlier task distributions are forgotten after later training. Banyan is a benchmark for studying when controlled task diversity produces transferable learning, when that transfer persists, and where it falls short of proper continual learning.

持续学习强化学习任务多样性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。