arXiv:2607.16895cs.LGcs.SY2026-07

安全控制器何时能适应?关键在于事前信息是否充足。

When Can Safe Controllers Adapt? Information before Commitment

论文配图:When Can Safe Controllers Adapt? Information before Commitment
图 1 · 摘自论文原文
  • 用因果归约分析安全控制的适应能力边界
  • 预承诺信息有限时,性能差距无法完全消除
  • 适用于需保证安全性的在线学习场景

安全自适应控制是在学习轨迹上保证安全的前提下进行在线调整。控制器可使用任意因果、依赖历史的规则,在不同环境中动态变化,但其安全保证必须在所有初始可能模型下统一成立。性能以知晓真实模型的安全预言机为基准衡量。多数有限时间分析假设一致安全闭环具有持续激励性,使数据能区分所有需不同控制决策的模型。在此前提下,可行性已确定,仅剩速率问题。本文提出新视角:安全约束本身是否允许此类有信息量的实验?只要替代模型仍可能成立,控制器就必须为其保留安全延续路径。首次导致该延续终止的动作称为‘承诺’。机会安全仅允许在替代模型下罕见事件发生时承诺,且证据必须在动作前抵达——动作产生的观测太晚。我们定义‘预承诺信息’为承诺前可观测规律的KL散度。主结果为因果归约:承诺规则决定了(1)在替代模型下安全允许承诺的概率,(2)维持非承诺状态的成本,(3)决策时刻可获取的信息。预承诺信息有界意味着固定比例的预言机差距不可避免。若该差距为Ω(T),则每个一致安全策略均有线性遗憾。我们在带约束的线性系统与二次调节成本下建立此障碍,并在特殊情况下证明可恢复性,还为确定性线性高斯系统导出半定上界证书。

原文摘要 · Abstract (English)

Safe adaptive control is online adaptation under a safety guarantee on the learning trajectory itself. The controller may use any causal, history-dependent rule and act differently across environments as data arrive. Only its safety guarantee is uniform: the same rule must satisfy it under every initially plausible model. Performance is measured against a safe oracle that knows the realized model. Many finite-time analyses assume persistent excitation of the uniformly safe closed loop, so the data distinguish every pair of models requiring different control decisions. Under that assumption, feasibility is already settled; only the rate remains. We ask instead: Do the safety constraints permit such an informative experiment at all? While an alternative remains plausible, the controller must preserve a safe continuation under it. We call the first action that forecloses such a continuation commitment. Chance safety allows commitment only on an event rare under the alternative, and the evidence must arrive beforehand: the observation generated by the committing action is too late. We define precommitment information as the KL divergence between learner-visible laws stopped before commitment. Our main result is a causal reduction. The commitment rule determines (1) the probability that safety permits commitment under the alternative, (2) the target-side cost of remaining noncommittal, (3) and the information available when the decision is made. Bounded precommitment information therefore leaves a fixed fraction of the oracle gap unavoidable. If the gap is Ω(T), every uniformly safe policy has linear regret. We establish the obstruction in a constrained linear system with quadratic regulation cost. We also prove recovery in special cases and derive semidefinite upper certificates for deterministic linear-Gaussian systems.

安全控制在线学习因果推理自适应系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。