arXiv:2411.01184cs.AIcs.LO2024-11被引 7

用逻辑规则动态调整奖励,让多智能体高效协作完成复杂任务

Guiding Multi-agent Multi-task Reinforcement Learning by a Hierarchical Framework with Logical Reward Shaping

  • 用线性时序逻辑表达任务间约束关系,指导奖励生成
  • 在模拟环境中的多任务实验中,性能显著优于传统方法
  • 适合需要可解释决策的多智能体协同场景

多智能体分层强化学习(MAHRL)被广泛用于复杂大规模环境中的智能决策。然而,现有算法多依赖传统奖励函数,仅适用于单任务。本文提出一种基于逻辑奖励塑造(LRS)的多智能体协作算法,利用线性时序逻辑(LTL)表达复杂任务中子任务间的逻辑关系,并基于设计的奖励结构评估LTL子公式的满足情况,使智能体通过遵循LTL表达式有效完成任务,提升决策的可解释性与可信度。为增强多智能体间协调,引入价值迭代技术评估各智能体行为,并据此构建协调奖励函数,使智能体可通过经验学习判断自身状态并完成剩余子任务。在类似Minecraft的环境中对多种任务类型进行实验,结果表明该算法在多任务学习中显著提升了多智能体性能。

原文摘要 · Abstract (English)

Multi-agent hierarchical reinforcement learning (MAHRL) has been studied as an effective means to solve intelligent decision problems in complex and large-scale environments. However, most current MAHRL algorithms follow the traditional way of using reward functions in reinforcement learning, which limits their use to a single task. This study aims to design a multi-agent cooperative algorithm with logic reward shaping (LRS), which uses a more flexible way of setting the rewards, allowing for the effective completion of multi-tasks. LRS uses Linear Temporal Logic (LTL) to express the internal logic relation of subtasks within a complex task. Then, it evaluates whether the subformulae of the LTL expressions are satisfied based on a designed reward structure. This helps agents to learn to effectively complete tasks by adhering to the LTL expressions, thus enhancing the interpretability and credibility of their decisions. To enhance coordination and cooperation among multiple agents, a value iteration technique is designed to evaluate the actions taken by each agent. Based on this evaluation, a reward function is shaped for coordination, which enables each agent to evaluate its status and complete the remaining subtasks through experiential learning. Experiments have been conducted on various types of tasks in the Minecraft-like environment. The results demonstrate that the proposed algorithm can improve the performance of multi-agents when learning to complete multi-tasks.

多智能体强化学习逻辑推理任务协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。