用量子强化学习优化智能表面,提升无线安全通信效率。
Securing SIM-Assisted Wireless Networks via Quantum Reinforcement Learning
- 将量子电路嵌入策略网络,增强高维连续动作空间的探索能力。
- 在信道信息不全时,秘密速率比传统深度强化学习高约15%,收敛快30%。
- 适合研究物理层安全与量子机器学习交叉方向的学者参考。
堆叠式智能超表面(SIM)作为一种新兴的波域技术,通过多层可编程结构实现对电磁信号的多级调控。尽管SIM提供了前所未有的自由度以增强物理层安全性,但其庞大的元原子数量导致优化空间维度极高且强耦合,传统设计方法效率低下且难以扩展。此外,现有深度强化学习(DRL)在动态无线环境中因对被动窃听者信道状态信息掌握不全,常出现收敛慢和性能下降问题。为此,本文提出一种混合量子近端策略优化(QPPO)框架,联合优化发射功率分配与SIM相位偏移,在功率和服务质量约束下最大化平均秘密速率。具体而言,将参数化量子电路嵌入演员网络,形成经典-量子混合策略架构,显著提升高维连续动作空间中的策略表达能力和探索效率。大量仿真实验表明,所提Q-PPO方案持续优于现有DRL基线,在信道信息不完全条件下,秘密速率提升约15%,收敛速度加快30%。这些结果确立了QPPO作为SIM赋能安全无线网络的强大优化范式。
原文摘要 · Abstract (English)
Stacked intelligent metasurfaces (SIMs) have recently emerged as a powerful wave-domain technology that enables multi-stage manipulation of electromagnetic signals through multilayer programmable architectures. While SIMs offer unprecedented degrees of freedom for enhancing physical-layer security, their extremely large number of meta-atoms leads to a high-dimensional and strongly coupled optimization space, making conventional design approaches inefficient and difficult to scale. Moreover, existing deep reinforcement learning (DRL) techniques suffer from slow convergence and performance degradation in dynamic wireless environments with imperfect knowledge of passive eavesdroppers. To address these challenges, we propose a hybrid quantum proximal policy optimization (QPPO) framework for SIM-assisted secure communications that jointly optimizes transmit power allocation and SIM phase shifts to maximize the average secrecy rate under power and quality-of-service constraints. Specifically, a parameterized quantum circuit is embedded into the actor network, forming a hybrid classical-quantum policy architecture that enhances policy representation capability and exploration efficiency in high-dimensional continuous action spaces. Extensive simulations demonstrate that the proposed Q-PPO scheme consistently outperforms DRL baselines, achieving approximately 15% higher secrecy rates and 30% faster convergence under imperfect eavesdropper channel state information. These results establish Q-PPO as a powerful optimization paradigm for SIM-enabled secure wireless networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。