arXiv:2503.03885cs.LG2025-03被引 2

让智能体在不协调情况下可靠协作,且保证安全行为。

Seldonian Reinforcement Learning for Ad Hoc Teamwork

  • 基于塞爾多尼优化,从数据中直接学习带安全保证的策略。
  • 在无需额外训练的情况下,比主流方法更高效地找到可靠策略。
  • 适合无人预先协同的多智能体场景,尤其安全敏感应用。

大多数离线强化学习算法虽能返回最优策略,但无法对期望行为提供统计保障,这在安全关键型多智能体场景中可能引发可靠性问题,例如智能体与新队友(甚至人类)需协作达成目标而不造成伤害。本文提出一种受塞爾多尼优化启发的新离线强化学习方法,可生成性能良好且在预定义理想行为上具有统计保障的策略。特别针对自适应团队协作(Ad Hoc Teamwork)场景——即智能体需与未事先协调的新队友合作。该方法仅需一个预收集数据集、一组候选策略及对其他智能体策略形式的规范说明,无需额外交互、训练或对策略类型与结构做假设。我们在自适应团队协作任务中测试该算法,结果表明其始终能发现可靠策略,并在样本效率上优于标准机器学习基线。

原文摘要 · Abstract (English)

Most offline RL algorithms return optimal policies but do not provide statistical guarantees on desirable behaviors. This could generate reliability issues in safety-critical applications, such as in some multiagent domains where agents, and possibly humans, need to interact to reach their goals without harming each other. In this work, we propose a novel offline RL approach, inspired by Seldonian optimization, which returns policies with good performance and statistically guaranteed properties with respect to predefined desirable behaviors. In particular, our focus is on Ad Hoc Teamwork settings, where agents must collaborate with new teammates without prior coordination. Our method requires only a pre-collected dataset, a set of candidate policies for our agent, and a specification about the possible policies followed by the other players -- it does not require further interactions, training, or assumptions on the type and architecture of the policies. We test our algorithm in Ad Hoc Teamwork problems and show that it consistently finds reliable policies while improving sample efficiency with respect to standard ML baselines.

强化学习多智能体安全约束离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。