arXiv:2604.24894cs.ROcs.CV2026-04被引 6

用视觉特征实现安全的实时控制,兼顾不确定性与系统约束。

VISION-SLS: Safe Perception-Based Control from Learned Visual Representations via System Level Synthesis

论文配图:VISION-SLS: Safe Perception-Based Control from Learned Visual Representations via System Level Synthesis
图 1 · 摘自论文原文
  • 从高分辨率图像中学习低维观测映射,带状态依赖误差界。
  • 在4D车、10D四旋翼等任务中实现安全探索,降低不确定性并满足约束。
  • 适合需要视觉感知与安全保障的机器人控制场景。

我们提出VISION-SLS,一种基于高分辨率RGB图像的非线性输出反馈控制方法,可在校准的不确定性边界下,应对部分可观测性、传感器噪声和非线性动态,仍保证约束满足。为实现可扩展性同时保留保证,我们提出:(i) 从预训练视觉特征中学习的低维观测映射,具有状态依赖误差界;(ii) 通过系统层级合成(SLS)优化的因果仿射时变输出反馈策略。我们开发了一种新型可扩展求解器,结合序列凸规划与高效Riccati递推,求解非凸问题。在两个模拟视觉运动任务(4D汽车、10D四旋翼,图像≥512×512像素)及一个59维人类机器人任务(部分可观测)上,该方法实现了安全的信息采集行为,减少不确定性,同时以经验校准的误差界保证约束满足。硬件验证显示,基于车载图像安全控制地面车辆,相比基线在安全率和求解时间上表现更优。结果表明,学习的视觉抽象结合高效求解器,使基于SLS的安全视觉运动输出反馈在大规模场景中成为可能。代码见https://github.com/trustworthyrobotics/VISION-SLS。

原文摘要 · Abstract (English)

We propose VISION-SLS, a method for nonlinear output-feedback control from high-resolution RGB images which provides robust constraint satisfaction guarantees under calibrated uncertainty bounds despite partial observability, sensor noise, and nonlinear dynamics. To enable scalability while retaining guarantees, we propose: (i) a learned low-dimensional observation map from pretrained visual features with state-dependent error bounds, and (ii) a causal affine time-varying output-feedback policy optimized via System Level Synthesis (SLS). We develop a scalable, novel solver for the resulting nonconvex program that leverages sequential convex programming coupled with efficient Riccati recursions. On two simulated visuomotor tasks (a 4D car and a 10D quadrotor) with >= 512 x 512 pixels and a 59D humanoid task with partial observability, our method enables safe, information-gathering behavior that reduces uncertainty while guaranteeing constraint satisfaction with empirically-calibrated error bounds. We also validate our method on hardware, safely controlling a ground vehicle from onboard images, outperforming baselines in safety rate and solve times. Together, these results show that learned visual abstractions coupled with an efficient solver make SLS-based safe visuomotor output-feedback practical at scale. The code implementation of our method is available at https://github.com/trustworthyrobotics/VISION-SLS.

视觉控制安全控制系统层级合成机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。