用分层图注意力模型提升对抗资源分配的决策效率
HGFormer: A Hierarchical Graph Transformer Framework for Two-Stage Colonel Blotto Games via Reinforcement Learning
- 构建分层图Transformer框架,融合结构先验与双智能体协同决策
- 在复杂动态博弈中实现更高资源利用效率和对抗收益
- 适合研究动态对抗策略、大规模资源分配的学者与工程师
两阶段红蓝军资源分配博弈是典型的对抗性资源分配问题,双方在两个阶段依次进行资源部署与动态调整。由于阶段间的序列依赖及图拓扑带来的复杂约束,传统方法难以获得全局最优策略。为此,我们提出一种名为HGformer的分层图Transformer框架,通过引入增强型图Transformer编码器(含结构偏置)与双智能体分层决策模型,实现了大规模对抗环境中的高效策略生成。同时设计逐层反馈强化学习算法,将低层决策的长期回报反馈至高层策略优化,弥合两阶段间的协调差距。实验表明,在复杂动态博弈场景中,相较现有层次决策或图神经网络方法,HGformer显著提升资源分配效率与对抗收益,整体表现更优。
原文摘要 · Abstract (English)
Two-stage Colonel Blotto game represents a typical adversarial resource allocation problem, in which two opposing agents sequentially allocate resources in a network topology across two phases: an initial resource deployment followed by multiple rounds of dynamic reallocation adjustments. The sequential dependency between game stages and the complex constraints imposed by the graph topology make it difficult for traditional approaches to attain a globally optimal strategy. To address these challenges, we propose a hierarchical graph Transformer framework called HGformer. By incorporating an enhanced graph Transformer encoder with structural biases and a two-agent hierarchical decision model, our approach enables efficient policy generation in large-scale adversarial environments. Moreover, we design a layer-by-layer feedback reinforcement learning algorithm that feeds the long-term returns from lower-level decisions back into the optimization of the higher-level strategy, thus bridging the coordination gap between the two decision-making stages. Experimental results demonstrate that, compared to existing hierarchical decision-making or graph neural network methods, HGformer significantly improves resource allocation efficiency and adversarial payoff, achieving superior overall performance in complex dynamic game scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。