arXiv:2606.16331cs.LG2026-06

用扩散模型提升无人机网络的节能与公平性调控

Diffusion Offline Reinforcement Learning for Fair and Energy-Efficient UAV-Assisted Wireless Networks

论文配图:Diffusion Offline Reinforcement Learning for Fair and Energy-Efficient UAV-Assisted Wireless Networks
图 1 · 摘自论文原文
  • 结合扩散模型与离线强化学习,生成更鲁棒的飞行与调度策略
  • 数据量少时仍保持稳定收敛,能耗降低35%以上,吞吐量显著提升
  • 适合6G无线网络中低数据场景下的智能控制,对算法鲁棒性要求高者适用

将生成式人工智能引入无线通信与信号处理系统,为未来6G网络提供智能化数据驱动决策。本文提出一种基于去噪扩散概率模型(DDPMs)增强的离线强化学习方法——扩散软演员-评论家(Diffusion-SAC),用于优化无人机(UAV)网络中的轨迹与调度控制。传统离线强化学习方法如保守Q学习(CQL)虽可从静态数据集中学习,但在低数据或动态环境下泛化能力差。为此,本工作融合CQL的稳健性与扩散模型的生成能力,实现表达性强且信号感知的策略学习,超越行为策略的限制。在无人机辅助无线网络中,该框架有效降低传输能耗,提升设备间公平性。仿真表明,与标准离线强化学习基线相比,Diffusion-SAC在数据有限条件下仍具更稳定的收敛性与更高奖励,数据效率提升,能耗降低超35%,吞吐量显著增加,展现出下一代无线控制系统中鲁棒策略学习的巨大潜力。

原文摘要 · Abstract (English)

The integration of generative artificial intelligence with wireless communication and signal processing systems has opened new avenues for intelligent, data-driven decision-making in future 6G networks. This work proposes a diffusion soft actor-critic (Diffusion-SAC) approach that leverages offline reinforcement learning (RL) enhanced by denoising diffusion probabilistic models (DDPMs) to optimize trajectory and scheduling control in unmanned aerial vehicle (UAV) networks. While offline RL methods, such as conservative Q-learning (CQL), can learn from static datasets, they often struggle to generalize in low-data or dynamic conditions. To address this, we combine the robustness of CQL with the generative power of diffusion models, enabling expressive and signal-aware policy learning that generalizes beyond behavior policies. Applied to a UAV-assisted wireless network, the proposed framework minimizes transmission energy and improves fairness among devices. Simulations show that Diffusion-SAC outperforms standard offline RL baselines, achieving more stable convergence and higher rewards even with limited datasets. The method enhances data efficiency, reduces energy consumption, and increases throughput by more than 35 % compared to existing algorithms, demonstrating its potential for robust policy learning in next-generation wireless control systems.

无人机网络强化学习能量效率扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。