用在线黎曼优化提升安全控制,让模型在复杂环境里更稳更快。
Safety-Critical Contextual Control via Online Riemannian Optimization with World Models

- 基于评分密度构建动作空间的黎曼几何,引导安全策略优化
- 安全性由条件曲率κ(ξ_t)决定,曲率越大越安全且收敛越快
- 新方法在环境变化后表现更好,适合动态场景的安全控制
现代世界模型越来越复杂,无法给出明确的动力学描述。本文研究安全关键的上下文控制问题:规划器仅能通过黑箱模拟器获取可行性样本,并根据上下文信号ξ_t优化任务目标。我们提出一种基于样本的惩罚预测控制(PPC)框架,基于在线黎曼优化,其中模拟器将可行性流形压缩为基于评分的密度估计̂p(u|ξ_t),赋予动作空间黎曼几何结构,指导规划器梯度下降。条件对数密度-̂ln̂p(̂̂̂·|ξ_t)的最小曲率κ(ξ_t)决定了收敛速度与安全裕度,替代了未知动力学的利普希茨常数。主要结果为上下文安全边界:真实可行性流形的距离受评分估计误差和依赖κ(ξ_t)的比值控制,两者均随上下文信息丰富而改善。动态导航任务的仿真表明,上下文式PPC显著优于边缘化与固定密度模型,且在环境变化后优势更大。
原文摘要 · Abstract (English)
Modern world models are becoming too complex to admit explicit dynamical descriptions. We study safety-critical contextual control, where a Planner must optimize a task objective using only feasibility samples from a black-box Simulator, conditioned on a context signal $ξ_t$. We develop a sample-based Penalized Predictive Control (PPC) framework grounded in online Riemannian optimization, in which the Simulator compresses the feasibility manifold into a score-based density $\hat{p}(u \mid ξ_t)$ that endows the action space with a Riemannian geometry guiding the Planner's gradient descent. The barrier curvature $κ(ξ_t)$, the minimum curvature of the conditional log-density $-\ln\hat{p}(\cdot\midξ_t)$, governs both convergence rate and safety margin, replacing the Lipschitz constant of the unknown dynamics. Our main result is a contextual safety bound showing that the distance from the true feasibility manifold is controlled by the score estimation error and a ratio that depends on $κ(ξ_t)$, both of which improve with richer context. Simulations on a dynamic navigation task confirm that contextual PPC substantially outperforms marginal and frozen density models, with the advantage growing after environment shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。