arXiv:2512.01502quant-phcs.AI2025-12

用形式化方法验证带噪声的量子强化学习策略是否安全。

Formal Verification of Noisy Quantum Reinforcement Learning Policies

  • 构建包含量子噪声的完整策略-环境交互模型,直接建模测量不确定性。
  • 在多个环境中验证:不同噪声会降低性能,但有时反而提升安全性。
  • 适合需提前验证安全性的量子强化学习部署场景。

量子强化学习(QRL)旨在利用量子效应设计比经典方法更有效的序列决策策略。然而,量子测量和硬件噪声(如比特翻转、相位翻转、去极化错误)会导致策略行为不可靠。现有方法缺乏系统性验证手段来评估训练后的QRL策略在特定噪声条件下的安全性。本文提出QVerifier,一种基于概率模型检测的形式化验证方法,可分析含噪与无噪的训练后QRL策略。QVerifier构建完整的策略-环境交互模型,将量子不确定性直接融入转移概率,并使用Storm模型检查器验证安全属性。实验表明,该方法能精确量化不同噪声模型对安全性的影响,揭示性能下降及噪声可能带来的益处。由于量子硬件成本高昂,部署前的严格验证至关重要。QVerifier适用于经典与量子计算的交汇点:当策略在匹配噪声条件下训练时,模型完全准确;若在真实硬件上训练,则构成理想化近似,因未知硬件噪声无法精确建模策略。

原文摘要 · Abstract (English)

Quantum reinforcement learning (QRL) aims to use quantum effects to create sequential decision-making policies that achieve tasks more effectively than their classical counterparts. However, QRL policies face uncertainty from quantum measurements and hardware noise, such as bit-flip, phase-flip, and depolarizing errors, which can lead to unsafe behavior. Existing work offers no systematic way to verify whether trained QRL policies meet safety requirements under specific noise conditions. We introduce QVerifier, a formal verification method that applies probabilistic model checking to analyze trained QRL policies with and without modeled quantum noise. QVerifier builds a complete model of the policy-environment interaction, incorporates quantum uncertainty directly into the transition probabilities, and then checks safety properties using the Storm model checker. Experiments across multiple QRL environments show that QVerifier precisely measures how different noise models influence safety, revealing both performance degradation and cases where noise can help. By enabling rigorous safety verification before deployment, QVerifier addresses a critical need: because access to quantum hardware is expensive, pre-deployment verification is essential for any safety-critical use of QRL. QVerifier targets a potential sweet spot between classical and quantum computation, where trained QRL policies could still be modeled classically for probabilistic model checking. When the policy was trained under matching noise conditions, this formal model is exact; when trained on physical hardware, it constitutes an idealized approximation, as unknown hardware noise prevents exact policy modeling.

量子强化学习形式化验证安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。