arXiv:2510.10214math.OCcs.AI2025-10

端到端学习可变度量,让鲁棒控制更精准高效。

Distributionally Robust Control with End-to-End Statistically Guaranteed Metric Learning

  • 将度量学习与控制任务闭环融合,动态调整不确定性集
  • 在数值和库存控制任务中性能超越现有方法
  • 理论保证有限样本下统计可靠性,适合高安全要求场景

Wasserstein分布鲁棒控制(DRC)是处理随机动力系统不确定性的新范式。然而,现有方法先统一构建模糊集,再逐次用于控制设计,导致模糊集构造与控制目标结构错配,产生保守且性能不佳的控制策略。为此,本文提出一种端到端的有限时域Wasserstein DRC框架,将各向异性Wasserstein度量学习与下游控制任务联合优化,使模糊集沿性能关键方向自适应调整,提升控制有效性。该框架为双层规划:内层描述受控系统演化,外层通过多初始条件下的控制性能反馈优化度量。我们设计了一种针对双层结构的随机增广拉格朗日算法,理论上证明了所学模糊集在新型半径调节机制下保持统计有限样本保证,并建立了双层问题的适定性(关于可学习度量的连续性)。进一步证明算法收敛至外层问题的驻点,且收敛速率具有非渐近统计一致性。实验在数值和库存控制任务中验证,该框架在闭环性能与鲁棒性上均优于当前最优方法。

原文摘要 · Abstract (English)

Wasserstein distributionally robust control (DRC) recently emerges as a principled paradigm for handling uncertainty in stochastic dynamical systems. However, it constructs data-driven ambiguity sets via uniform distribution shifts before sequentially incorporating them into downstream control synthesis. This segregation between ambiguity set construction and control objectives inherently introduces a structural misalignment, which undesirably leads to conservative control policies with sub-optimal performance. To address this limitation, we propose a novel end-to-end finite-horizon Wasserstein DRC framework that integrates the learning of anisotropic Wasserstein metrics with downstream control tasks in a closed-loop manner, thus enabling ambiguity sets to be systematically adjusted along performance-critical directions and yielding more effective control policies. This framework is formulated as a bilevel program: the inner level characterizes dynamical system evolution under DRC, while the outer level refines the anisotropic metric leveraging control-performance feedback across a range of initial conditions. To solve this program efficiently, we develop a stochastic augmented Lagrangian algorithm tailored to the bilevel structure. Theoretically, we prove that the learned ambiguity sets preserve statistical finite-sample guarantees under a novel radius adjustment mechanism, and we establish the well-posedness of the bilevel formulation by demonstrating its continuity with respect to the learnable metric. Furthermore, we show that the algorithm converges to stationary points of the outer level problem, which are statistically consistent with the optimal metric at a non-asymptotic convergence rate. Experiments on both numerical and inventory control tasks verify that the proposed framework achieves superior closed-loop performance and robustness compared against state-of-the-art methods.

鲁棒控制分布鲁棒度量学习双层优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。