提出PDTS方法,让智能体在随机环境中更快更稳地适应新任务。
Fast and Robust: Task Sampling with Posterior and Diversity Synergies for Adaptive Decision-Makers in Randomized Environments
- 基于后验概率与任务多样性协同设计采样策略
- 显著提升零样本和少样本下的鲁棒适应能力
- 适合需要快速适应复杂环境的强化学习应用
序列决策中的任务鲁棒自适应是长期追求的目标。现有风险规避策略(如条件风险价值)常用于领域随机化或元强化学习中以优先优化困难任务,但需大量耗时评估。为解决效率问题,研究提出稳健主动任务采样以训练自适应策略,使用风险预测模型替代实际策略评估。本文将稳健主动任务采样建模为马尔可夫决策过程,给出理论与实践洞见,并构建风险规避场景下的鲁棒性概念。重要的是,提出一种易实现的方法——后验与多样性协同任务采样(PDTS),支持快速且稳健的序列决策。大量实验表明,PDTS充分发挥了稳健主动任务采样的潜力,在挑战性任务中显著提升零样本与少样本适应鲁棒性,甚至在某些场景下加速学习过程。
原文摘要 · Abstract (English)
Task robust adaptation is a long-standing pursuit in sequential decision-making. Some risk-averse strategies, e.g., the conditional value-at-risk principle, are incorporated in domain randomization or meta reinforcement learning to prioritize difficult tasks in optimization, which demand costly intensive evaluations. The efficiency issue prompts the development of robust active task sampling to train adaptive policies, where risk-predictive models are used to surrogate policy evaluation. This work characterizes the optimization pipeline of robust active task sampling as a Markov decision process, posits theoretical and practical insights, and constitutes robustness concepts in risk-averse scenarios. Importantly, we propose an easy-to-implement method, referred to as Posterior and Diversity Synergized Task Sampling (PDTS), to accommodate fast and robust sequential decision-making. Extensive experiments show that PDTS unlocks the potential of robust active task sampling, significantly improves the zero-shot and few-shot adaptation robustness in challenging tasks, and even accelerates the learning process under certain scenarios. Our project website is at https://thu-rllab.github.io/PDTS_project_page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。