让视觉控制的机器人预测安全概率,还能自校准信心。
How Safe Will I Be Given What I Saw? Calibrated Prediction of Safety Chances for Image-Controlled Autonomy
- 用变分自编码器和循环预测器从图像序列推未来状态轨迹。
- 在分布偏移下仍保持低误报率,长时预测准确率提升显著。
- 适合自动驾驶、机器人等对安全性要求高的视觉控制场景。
依赖深度神经网络控制器的自主机器人在部分可观测性和分布偏移下面临重大安全预测挑战。传统基于模型的方法难以扩展且需低维状态模型,而无模型方法常缺乏可靠性保障。本文提出一种端到端视觉控制系统中的校准安全预测框架,无需访问状态转移或观测模型。基于世界模型,利用变分自编码器与递归预测器从原始图像序列预测未来潜在轨迹,并评估满足安全属性的概率。区分单一式与复合式预测流程,引入校准机制量化预测置信度。在高维观测的长时预测中,因累积误差与环境变化,预测输入可能偏离训练分布,导致风险估计失准。为此,引入无监督域适应(UDA)以增强预测分布偏移下的安全评估鲁棒性,无需人工标注。理论保证预测校准性,支持长时程实际评估。三个基准测试结果表明,配备UDA的安全评估器在分布偏移下保持高精度与显著更低的误报率;世界模型驱动的复合预测器优于单一结构,且我们的共形校准提供可靠的统计界。
原文摘要 · Abstract (English)
Autonomous robots that rely on deep neural network controllers pose critical challenges for safety prediction, especially under partial observability and distribution shift. Traditional model-based verification techniques are limited in scalability and require access to low-dimensional state models, while model-free methods often lack reliability guarantees. This paper addresses these limitations by introducing a framework for calibrated safety prediction in end-to-end vision-controlled systems, where neither the state-transition model nor the observation model is accessible. Building on the foundation of world models, we leverage variational autoencoders and recurrent predictors to forecast future latent trajectories from raw image sequences and estimate the probability of satisfying safety properties. We distinguish between monolithic and composite prediction pipelines and introduce a calibration mechanism to quantify prediction confidence. In long-horizon predictions from high-dimensional observations, the forecasted inputs to the safety evaluator can deviate significantly from the training distribution due to compounding prediction errors and changing environmental conditions, leading to miscalibrated risk estimates. To address this, we incorporate unsupervised domain adaptation to ensure robustness of safety evaluation under distribution shift in predictions without requiring manual labels. Our formulation provides theoretical calibration guarantees and supports practical evaluation across long prediction horizons. Experimental results on three benchmarks show that our UDA-equipped evaluators maintain high accuracy and substantially lower false positive rates under distribution shift. Similarly, world model-based composite predictors outperform their monolithic counterparts on long-horizon tasks, and our conformal calibration provides reliable statistical bounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。