用概率保证方法为随机控制策略提供闭环轨迹的严格安全边界。
PAC-Bayesian Certificates for Quadratic Closed-Loop Control
- 基于系统级综合参数化,将二次控制损失转化为可认证形式
- 在低数据场景下提升预测性能并降低闭环敏感度,实测效果显著
- 适合对安全性要求高的学习型控制系统设计者使用
PAC-Bayesian界为数据依赖的随机预测器提供了有限样本保障,但将其应用于基于学习的控制却面临挑战,因为自然目标是二次轨迹代价。此类损失无界、非Lipschitz,且导致响应相关的Chernoff项。本文采用系统级综合(SLS)参数化,直接暴露线性系统的闭环轨迹映射,使二次控制损失可显式认证。针对任意协方差的高斯扰动轨迹,推导出精确的一侧高斯变换和可通过闭环灵敏度量表示的可处理二次上界。还提出一种后验局部代理,适用于点态闭环响应证书不可用或支持相关容许性问题的情况。尽管PAC-Bayes认证的是非退化后验,但SLS损失的凸二次形式仍可将证书转移至后验均值响应。本文给出一个确定性均值响应部署结果,特别适合控制应用,同时在界中保留随机后验。此外,提供一种数据驱动的该部署的界,摆脱了对先验知识的依赖。最小化该界自然导出一种从数据中学习控制选择的算法。双积分器系统的数值实验表明,该算法在低数据条件下表现如感知灵敏度的有限样本正则化器,有效降低外推代价并减少闭环敏感度。
原文摘要 · Abstract (English)
PAC-Bayesian bounds provide finite-sample guarantees for data-dependent randomized predictors, but applying them to learning-based control is difficult because the natural objective is a quadratic trajectory cost. Such losses are unbounded, non-Lipschitz , and lead to response-dependent Chernoff terms. We employ System Level Synthesis parameterization, which exposes the closed-loop trajectory map of a linear system directly and makes the quadratic control loss amenable to explicit certification. Moreover, we provide a set of PAC-Bayes-Chernoff certificates for posterior distributions over feasible closed-loop responses. For Gaussian disturbance trajectories with arbitrary covariance, we derive an exact one-sided Gaussian transform and a tractable quadratic upper bound expressed through closed-loop sensitivity quantities. We also derive a posterior-localized surrogate for settings where pointwise closed-loop response certificates are unavailable or have support related admissibility issues. Although PAC-Bayes certifies a non-degenerate posterior, the convex quadratic form of the SLS loss transfers the certificate to the posterior mean response. We present a deterministic mean response deployment result that is particularly suitable for control while retaining the stochastic posterior in the bound. Additionally, we provide a data-driven bound for this deployment, transitioning away from an oracle bound. Minimizing this bound naturally results in a learning algorithm for control selection from data. Numerical experiments on a double integrator show that the algorithm acts as a sensitivity-aware finite-sample regularizer, improving held-out cost and reducing closed-loop sensitivity in the low-data regime
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。