arXiv:2608.07265math.OCcs.LG2026-08

让控制模型在有限样本下保持几何精度,避免状态坍缩。

Finite-Sample Metric Non-Collapse for Geometrically Supervised Latent World Models in Control

  • 用几何监督训练编码器,分离方向与距离信息以防止状态折叠。
  • 理论证明优化后模型在高概率下保持度量一致性和动态近似性。
  • 适合关注模型可靠性与规划安全性的强化学习研究者。

本文建立了非线性确定性系统几何监督潜空间模型的有限样本学习-控制理论。几何监督仅用于训练:通过具有独立验证度量和方向误差界的状态、本体感知或状态估计提供可观测状态距离与切向方向,部署时仍为观测和动作条件。提出仅编码器的局部-全局度量铰链机制,强制方向分辨与状态分离。在可观测因子、覆盖性、有限容量逼近及统一C^{1,1}假设下,可计算的一侧正则化范式具有强选择性:以高概率,每个近似经验最小化解同时满足逐点共利普希茨和均匀近似半共轭于受控动力学。逼近、采样与优化误差保持显式且分离。范数约束张量积B样条类构造性实现逼近假设,将均方残差控制转化为一致界所需的插值指数是紧的。模块化确定性推论将所学证书转移至轨迹、有限时域代价、学习代价头和优化器保证;经验证的有限网结果实现更精确的模型特定认证。受控实验分离了坍缩与折叠现象,量化了分析证书的冗余,并展示了恢复度量分辨率的控制优势。核心贡献是:从近似经验优化到度量忠实性、一致受控动力学与可靠规划的完整有限样本推导。

原文摘要 · Abstract (English)

We establish a finite-sample learning-to-control theory for geometrically supervised latent models of nonlinear deterministic systems. Geometric supervision is used only during training: simulator state, proprioception, or state estimates with independently validated metric and directional error bounds supply observable-state distances and tangent directions, while deployment remains observation- and action-conditioned. We introduce an encoder-only local--global metric hinge that enforces directional resolution and separated-state discrimination. Under regular observable-factor, coverage, finite-capacity approximation, and uniform $C^{1,1}$ hypotheses, a computable one-sided regularization regime has a strong selection property: with high probability, every approximate empirical minimizer is simultaneously pointwise co-Lipschitz and uniformly approximately semiconjugate to the controlled dynamics. Approximation, sampling, and optimization errors remain explicit and separate. Norm-constrained tensor-product B-spline classes constructively realize the approximation hypotheses, and the interpolation exponent converting mean residual control into a uniform bound is sharp. A modular deterministic corollary transfers the learned certificates to trajectory, finite-horizon cost, learned-cost-head, and optimizer guarantees, while a validated finite-net result enables sharper model-specific certification. Controlled experiments isolate collapse and folding, quantify the analytic certificate's reserve, and demonstrate the control benefit of restored metric resolution. The principal contribution is a complete finite-sample implication from approximate empirical optimization to metric faithfulness, uniform controlled dynamics, and reliable planning for the same learned model.

控制理论潜空间建模度量学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。