无需手动设定偏好,仅靠示范数据就能自动学习多目标强化学习策略。
An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement Learning
- 基于示范数据隐式推断目标偏好,无需人工设计参数
- 在未知安全阈值下仍能生成符合约束的安全策略
- 适用于缺乏先验知识的复杂多目标任务,尤其适合安全敏感场景
近年来,多目标强化学习(Multi-Objective Reinforcement Learning, MORL)取得显著进展,旨在通过引入各目标的偏好来平衡多个目标。然而,现有方法通常需要在部署时提供明确的目标偏好,而这些偏好高度依赖人类先验知识,通常需通过观察高性能示范行为获得。本文提出一种简单的离线适应框架,无需预设目标偏好,仅需若干示范数据即可隐式指示期望策略的偏好。此外,我们证明该框架可自然扩展至处理安全关键目标的约束问题,即使安全阈值未知,也能利用安全示范数据生成满足约束的策略。在离线多目标与安全任务上的实证结果表明,该框架能有效推断出与真实偏好一致且满足示范所隐含约束的策略。
原文摘要 · Abstract (English)
In recent years, significant progress has been made in multi-objective reinforcement learning (RL) research, which aims to balance multiple objectives by incorporating preferences for each objective. In most existing studies, specific preferences must be provided during deployment to indicate the desired policies explicitly. However, designing these preferences depends heavily on human prior knowledge, which is typically obtained through extensive observation of high-performing demonstrations with expected behaviors. In this work, we propose a simple yet effective offline adaptation framework for multi-objective RL problems without assuming handcrafted target preferences, but only given several demonstrations to implicitly indicate the preferences of expected policies. Additionally, we demonstrate that our framework can naturally be extended to meet constraints on safety-critical objectives by utilizing safe demonstrations, even when the safety thresholds are unknown. Empirical results on offline multi-objective and safe tasks demonstrate the capability of our framework to infer policies that align with real preferences while meeting the constraints implied by the provided demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。