让擅长的任務指導新手任務,加速多任務強化學習。
Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance
- 用引導策略從其他任務中選擇最佳行為策略,生成更優訓練軌跡。
- 在操控與移動基準上提升性能,顯著加快技能習得速度。
- 適配現有共享參數方法,適合多任務強化學習研究者使用。
多任務強化學習旨在高效利用不同任務間的共性信息,實現多任務同時學習。現有方法主要依賴精心設計的網絡結構或專門的優化過程進行參數共享,但忽略了另一種直接且補充性的跨任務相似性利用方式:已掌握某些技能的任務控制策略可為未掌握技能的任務提供明確指導,以加速技能習得。為此,我們提出一種新框架——跨任務策略引導(CTPG),為每個任務訓練一個引導策略,從所有任務的控制策略中選擇與環境交互的最佳行為策略,從而生成更優的訓練軌跡。此外,我們設計了兩種閘控機制:一種過濾對引導無益的控制策略,另一種阻斷無需引導的任務。CTPG是一種通用框架,可與現有參數共享方法兼容。實驗表明,將CTPG與這些方法結合,在操控與運動基準上均顯著提升性能。
原文摘要 · Abstract (English)
Multi-task reinforcement learning endeavors to efficiently leverage shared information across various tasks, facilitating the simultaneous learning of multiple tasks. Existing approaches primarily focus on parameter sharing with carefully designed network structures or tailored optimization procedures. However, they overlook a direct and complementary way to exploit cross-task similarities: the control policies of tasks already proficient in some skills can provide explicit guidance for unmastered tasks to accelerate skills acquisition. To this end, we present a novel framework called Cross-Task Policy Guidance (CTPG), which trains a guide policy for each task to select the behavior policy interacting with the environment from all tasks' control policies, generating better training trajectories. In addition, we propose two gating mechanisms to improve the learning efficiency of CTPG: one gate filters out control policies that are not beneficial for guidance, while the other gate blocks tasks that do not necessitate guidance. CTPG is a general framework adaptable to existing parameter sharing approaches. Empirical evaluations demonstrate that incorporating CTPG with these approaches significantly enhances performance in manipulation and locomotion benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。