用强化学习优化卫星通信链路配置,验证其在资源调度中的潜力。
Optimization of Link Configuration for Satellite Communication Using Reinforcement Learning
- 设计卫星转发器环境,用PPO算法进行链路配置优化。
- 模拟退火在静态问题中表现优于PPO算法。
- 证明强化学习在复杂资源优化中有应用前景,适合研究者参考。
卫星通信是现代互联世界的关键技术。随着硬件日益复杂,如何高效配置卫星转发器上的链路成为难题,该问题涉及众多参数与指标,合理利用有限的带宽和功率资源至关重要。尽管以往可用模拟退火等元启发式方法近似求解,但近期研究表明强化学习在优化任务中可达到甚至超越传统方法性能。然而,此前尚无针对卫星转发器链路配置的研究。为此,本文构建了转发器仿真环境,并在两项实验中对比了PPO算法与模拟退火的表现。结果表明,在该静态问题下,模拟退火优于PPO;但研究也凸显了强化学习在复杂优化中的潜力。
原文摘要 · Abstract (English)
Satellite communication is a key technology in our modern connected world. With increasingly complex hardware, one challenge is to efficiently configure links (connections) on a satellite transponder. Planning an optimal link configuration is extremely complex and depends on many parameters and metrics. The optimal use of the limited resources, bandwidth and power of the transponder is crucial. Such an optimization problem can be approximated using metaheuristic methods such as simulated annealing, but recent research results also show that reinforcement learning can achieve comparable or even better performance in optimization methods. However, there have not yet been any studies on link configuration on satellite transponders. In order to close this research gap, a transponder environment was developed as part of this work. For this environment, the performance of the reinforcement learning algorithm PPO was compared with the metaheuristic simulated annealing in two experiments. The results show that Simulated Annealing delivers better results for this static problem than the PPO algorithm, however, the research in turn also underlines the potential of reinforcement learning for optimization problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。