arXiv:2504.05045cs.LGcs.MA2025-04被引 2

用时空注意力约束奖励学习,提升多智能体任务分配的稳定性与效率。

Spatiotemporal Attention-Augmented Inverse Reinforcement Learning for Multi-Agent Task Allocation

  • 引入时空注意力结构,建模长时序依赖和智能体-任务关系。
  • 奖励推断采用低容量自适应线性变换,收敛更快且累积收益更高。
  • 适合复杂多智能体系统中需要稳定奖励学习的场景。

针对多智能体任务分配(MATA)中的对抗式逆强化学习(IRL)面临非平稳交互与高维协调挑战,现有方法在无约束奖励推断下常导致方差大、泛化差。本文提出一种基于注意力结构的对抗式IRL框架,通过时空表征学习约束奖励推断。方法采用多头自注意力(MHSA)捕捉长期时间依赖,图注意力网络(GAT)建模智能体-任务关系。将奖励推断建模为环境奖励的低容量、自适应线性变换,实现稳定可解释的引导。该框架解耦奖励推断与策略学习,并通过对抗方式优化奖励模型。在基准MATA场景上的实验表明,本方法在收敛速度、累积奖励与空间效率上优于代表性多智能体强化学习基线。结果验证了注意力引导、容量受限的奖励推断是复杂多智能体系统中稳定对抗IRL的有效机制。

原文摘要 · Abstract (English)

Adversarial inverse reinforcement learning (IRL) for multi-agent task allocation (MATA) is challenged by non-stationary interactions and high-dimensional coordination. Unconstrained reward inference in these settings often leads to high variance and poor generalization. We propose an attention-structured adversarial IRL framework that constrains reward inference via spatiotemporal representation learning. Our method employs multi-head self-attention (MHSA) for long-range temporal dependencies and graph attention networks (GAT) for agent-task relational structures. We formulate reward inference as a low-capacity, adaptive linear transformation of the environment reward, ensuring stable and interpretable guidance. This framework decouples reward inference from policy learning and optimizes the reward model adversarially. Experiments on benchmark MATA scenarios show that our approach outperforms representative MARL baselines in convergence speed, cumulative rewards, and spatial efficiency. Results demonstrate that attention-guided, capacity-constrained reward inference is a scalable and effective mechanism for stabilizing adversarial IRL in complex multi-agent systems.

多智能体逆强化学习注意力机制任务分配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。