arXiv:2501.11921cs.ITcs.AI2025-01被引 1

提出新算法提升智能通信系统调度效率,性能比现有方法高45%。

Goal-oriented Transmission Scheduling: Structure-guided DRL with a Unified Dual On-policy and Off-policy Approach

  • 基于状态结构特性设计混合强化学习算法
  • 性能提升45%,收敛速度加快40%
  • 适合大规模多设备无线系统调度场景

面向目标的通信以应用需求为核心,推动下一代智能无线系统发展。在多设备、多信道系统中,高维状态与动作空间带来调度难题。本文推导出目标导向调度最优解的关键结构特性,结合信息年龄(AoI)与信道状态,证明了最优状态值函数关于信道状态的单调性及关于AoI状态的渐近凸性,并得出最优策略对信道状态的单调性,完善了调度理论基础。据此提出结构引导的统一双策略深度强化学习(SUDO-DRL),融合在线策略稳定性与离线策略样本效率。通过新型结构评估框架,实现高效可扩展训练,在大规模系统中表现优越。数值结果表明,相比最先进方法,系统性能提升最高达45%,收敛时间减少40%;在传统离线策略失效、在线策略性能显著下降的大规模场景中仍保持高效,验证了其在目标导向通信中的可扩展性与有效性。

原文摘要 · Abstract (English)

Goal-oriented communications prioritize application-driven objectives over data accuracy, enabling intelligent next-generation wireless systems. Efficient scheduling in multi-device, multi-channel systems poses significant challenges due to high-dimensional state and action spaces. We address these challenges by deriving key structural properties of the optimal solution to the goal-oriented scheduling problem, incorporating Age of Information (AoI) and channel states. Specifically, we establish the monotonicity of the optimal state value function (a measure of long-term system performance) w.r.t. channel states and prove its asymptotic convexity w.r.t. AoI states. Additionally, we derive the monotonicity of the optimal policy w.r.t. channel states, advancing the theoretical framework for optimal scheduling. Leveraging these insights, we propose the structure-guided unified dual on-off policy DRL (SUDO-DRL), a hybrid algorithm that combines the stability of on-policy training with the sample efficiency of off-policy methods. Through a novel structural property evaluation framework, SUDO-DRL enables effective and scalable training, addressing the complexities of large-scale systems. Numerical results show SUDO-DRL improves system performance by up to 45% and reduces convergence time by 40% compared to state-of-the-art methods. It also effectively handles scheduling in much larger systems, where off-policy DRL fails and on-policy benchmarks exhibit significant performance loss, demonstrating its scalability and efficacy in goal-oriented communications.

强化学习无线通信调度优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。