用智能无人机网络优化灾后信息更新速度,让救援更及时。
MA-HEAD-Net: Adaptive Rule-Guided Multi-Agent DRL for AoI Minimization in UAV-Assisted Emergency Networks
- 设计自适应调度机制,动态调整数据传输节奏
- 实测比主流方法降低30%以上信息延迟
- 适合应急通信、无人机协同等实时系统研究者
灾后场景中,无人机对建立应急通信网络至关重要。为保障救援决策的时效性,需最小化信息年龄(AoI)。本文针对异构应急服务下的无人机通信,采用马尔可夫调制泊松过程建模突发数据到达,并结合有限块长理论刻画传输时长、包完成率与AoI演化间的耦合关系。提出嵌入小时间片的调度机制,支持自适应检查点间隔选择。将无人机轨迹控制、用户调度与检查点选择联合优化建模为多智能体决策问题,构建MA-HEAD-Net——一种融合通信规则先验的自适应多头深度强化学习框架。该框架通过自适应门控调节规则先验与学习策略在不同子任务中的贡献,采用多智能体近端策略优化联合训练。仿真表明,相比主流多智能体强化学习方法,MA-HEAD-Net显著提升策略生成效率,在动态场景下实现更低的平均信息年龄,优于基于学习和启发式的方法。
原文摘要 · Abstract (English)
In post-disaster scenarios, unmanned aerial vehicles (UAVs) are critical for establishing emergency communication networks. For time-critical rescue missions, information freshness is crucial because decisions based on outdated data may lead to ineffective control actions. This paper investigates age of information (AoI) minimization for UAV-assisted emergency communications with heterogeneous emergency services. We model bursty packet arrivals using a Markov-modulated Poisson process and adopt finite blocklength theory to capture the coupling among transmission duration, packet completion, and AoI evolution. To balance delay-tolerant long-packet transmission and urgent short-packet response, we propose a mini-slot-embedded scheduling mechanism with adaptive checkpoint-interval selection. We formulate the joint optimization of UAV trajectory control, user scheduling, and checkpoint-interval selection as a multi-agent decision problem, and develop MA-HEAD-Net, an adaptive rule-guided multi-agent deep reinforcement learning framework. MA-HEAD-Net incorporates communication-domain rule priors into a gated multi-head policy, where adaptive gates regulate the contributions of rule-prior and learned-policy logits for different subtasks. The policy and gating components are jointly optimized under multi-agent proximal policy optimization. Simulation results show that MA-HEAD-Net improves policy-formation efficiency compared with representative multi-agent deep reinforcement learning baselines and achieves lower AoI than both learning-based and heuristic methods in dynamic UAV-assisted emergency communication scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。