arXiv:2509.08654quant-phcs.AI2025-09

在量子网络动态退相干下,用信念状态实现鲁棒路由决策。

Robust Belief-State Policy Learning for Quantum Network Routing Under Decoherence and Time-Varying Conditions

  • 基于量子部分可观测马尔可夫决策过程与带可行性掩码的图神经网络
  • 在有限内存拓扑上提升高保真度吞吐量,降低低于阈值交付率
  • 适合需要资源约束下实时决策的量子网络控制场景

量子网络路由需在纠缠生成概率性、有限量子内存、退相干、操作不完美及经典反馈条件下进行在线决策,且控制器对物理状态知之不全。本文提出一种基于量子部分可观测马尔可夫决策过程(q-POMDP)和可行性掩码图神经网络(GNN)的鲁棒信念状态路由框架。模型采用原子微周期,确保每项操作在下一决策边界前完成,从而显式建模内存预留、配对实例库存、纯化消耗、交换结果、释放决策、队列服务及完成时间保真度。控制器维护对隐藏物理状态的古典信念,包括潜在环境条件,并据此评估可行动作与更新后验配对状态。为提升可扩展性,引入可行性分层原型、无标识符签名与角色感知动作匹配,在保留硬性资源约束的同时实现信息状态间价值迁移。随后通过自适应信任规则将缓存q-POMDP规划器与角色感知GNN策略融合,并为未见可行性签名提供安全回退。理论证明涵盖可行性、价值近似、策略性能、鲁棒性、遗憾与学习方差。仿真显示,该混合控制器在有限内存量子网络拓扑上优于纯规划、启发式、纯化感知及学习基线,显著提升高保真度吞吐量,减少低于阈值交付,且在线决策成本更低。

原文摘要 · Abstract (English)

Quantum network routing requires online decisions under probabilistic entanglement generation, finite quantum memories, decoherence, imperfect operations, and classical feedback, while the controller has incomplete knowledge of the physical state. This paper develops a robust belief-state routing framework based on a quantum partially observable Markov decision process (q-POMDP) and a feasibility-masked graph neural network (GNN). The model uses atomic micro-epochs in which each selected operation completes before the next decision boundary. This enables explicit accounting of memory reservations, pair-instance inventories, purification consumption, swapping outcomes, release decisions, queue service, and completion-time delivery fidelity. The controller maintains a classical belief over hidden physical states, including latent environmental conditions, and uses this belief to evaluate feasible actions and update posterior pair states. To make planning scalable, we introduce feasibility-stratified prototypes, identifier-free signatures, and role-aware action matching, which preserve hard resource constraints while enabling value transfer across structurally similar information states. A cached q-POMDP planner is then fused with a role-aware GNN policy through an adaptive trust rule, with a safe fallback for previously unseen feasibility signatures. We provide theoretical guarantees on feasibility, value approximation, policy performance, robustness, regret, and learning variance. Simulations over finite-memory quantum-network topologies show that the proposed hybrid controller improves high-fidelity goodput, reduces below-threshold deliveries, and maintains lower online decision cost than planner-only control, while outperforming heuristic, purification-aware, and learning-based baselines.

量子网络信念状态强化学习资源约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。