arXiv:2606.09778quant-phcs.AI2026-06

提出可衡量安全责任归属的量子控制方法,让政策自己扛安全

Who Earns the Safety? Intervention-Aware Quantum Predictive Control with Safety Attribution

论文配图:Who Earns the Safety? Intervention-Aware Quantum Predictive Control with Safety Attribution
图 1 · 摘自论文原文
  • 用量子电路+干预预算训练策略,减少对安全层依赖
  • 实验显示量子策略违规率和依赖度显著降低(均p<10^-4)
  • 适合关注安全责任归因与量子控制的学者

为确保运行时约束满足,越来越多学习型控制器下游加入硬性安全过滤器。但未违反约束的控制器可能并未真正学会安全:过滤器可能默默修复上游策略缺陷,使事后成功率反映的是过滤器而非策略本身。我们主张安全策略学习应追问‘谁赚了安全’——是策略还是保护层,并使其可测量。提出干预感知变分量子可微预测控制(IA-VQC-DPC),(i) 在原始-对偶干预预算下训练紧凑变分量子电路(VQC)策略,惩罚对可微控制屏障函数(CBF)投影的依赖;(ii) 通过安全归因协议评估,将轨迹修正分解为CBF项与部署运行时保护项,并进行无保护评估。在闭环、高保真BOPTEST建筑控制模拟器上(5个随机种子,每方法60次试验),干预感知训练显著降低量子策略的原始预过滤违规率和整体安全层依赖度(均p < 10^-4),且无显著能耗退化;在约400参数预算下,量子策略比对应经典策略更安全、更舒适。无保护评估确认改进来自策略层面,并暴露一个重要负面结果:学习到的可微能量头仅在搭配分布感知运行时保护时才安全。该归因协议适用于超越量子策略与建筑场景的广泛情形。

原文摘要 · Abstract (English)

Hard safety filters are increasingly placed downstream of learned controllers to guarantee constraint satisfaction at run time. Yet a filtered controller that never violates a constraint may still have learned nothing about safety: the filter can silently repair an incompetent upstream policy, so that post-filter success measures the filter, not the policy. We argue that safe policy learning should ask who earns the safety - the policy or its protective layers - and we make this question measurable. We introduce Intervention-Aware Variational Quantum Differentiable Predictive Control (IA-VQC-DPC), which (i) trains a compact variational quantum circuit (VQC) policy under a primal-dual intervention budget that penalizes reliance on a differentiable Control-Barrier-Function (CBF) projection, and (ii) is evaluated with a safety-attribution protocol that decomposes the executed-trajectory correction into a CBF term and a deployment runtime-guard term, and stress-tests the policy with guard-off evaluation. On closed-loop, high-fidelity BOPTEST building-control emulators (5 seeds, 60 episodes per method), intervention-aware training significantly lowers the quantum policy's raw pre-filter violation and total safety-layer reliance (both p < 10^-4) with no significant energy regression; at an equal approximately 400-parameter budget the quantum policy is significantly safer and more comfortable than a matched classical policy. Guard-off evaluation confirms the improvement is policy-level and exposes a valuable negative result: a learned differentiable energy head is only safe when paired with a distribution-aware runtime guard. The attribution protocol is general beyond quantum policies and buildings.

量子控制安全归因强化学习建筑节能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。