为多任务强化学习提供新任务的高置信度性能保证
Probabilistic Performance Guarantees for Multi-Task Reinforcement Learning
- 基于有限采样任务和轨迹构建任务级泛化边界
- 在真实样本量下对未见任务给出可信赖的性能下界
- 适合安全关键场景中对策略可靠性有要求的研究者
多任务强化学习训练通用策略以执行多种任务。尽管近年进展显著,现有方法很少提供正式的性能保证,而这类保证在安全关键场景部署策略时至关重要。本文提出一种计算多任务策略在训练中未见任务上高置信度性能保证的方法。具体而言,我们引入一个新的泛化界,该界由(i)从有限次滚动采样得到的各任务下置信界与(ii)从有限采样任务中获得的任务级泛化性组合而成,从而对来自相同任意未知分布的新任务提供高置信度保证。在多个主流多任务强化学习方法上,我们验证了该保证在实际样本量下理论成立且具有信息量。
原文摘要 · Abstract (English)
Multi-task reinforcement learning trains generalist policies that can execute multiple tasks. While recent years have seen significant progress, existing approaches rarely provide formal performance guarantees, which are indispensable when deploying policies in safety-critical settings. We present an approach for computing high-confidence guarantees on the performance of a multi-task policy on tasks not seen during training. Concretely, we introduce a new generalisation bound that composes (i) per-task lower confidence bounds from finitely many rollouts with (ii) task-level generalisation from finitely many sampled tasks, yielding a high-confidence guarantee for new tasks drawn from the same arbitrary and unknown distribution. Across state-of-the-art multi-task RL methods, we show that the guarantees are theoretically sound and informative at realistic sample sizes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。