用视觉数据实时验证复杂安全规则,无需重新训练。
Vision-Based Runtime Monitoring under Varying Specifications using Semantic Latent Representations

- 以语义基向量为最小预测目标,实现可复用的实时监控。
- 在长时序场景下,语义基监控比滚动预测更紧致,最多提升4倍。
- 适用于自动驾驶等对安全性要求高的真实场景,有严格统计保障。
本文研究在部分可观测条件下,基于视觉观测对过去时间信号时序逻辑(ptSTL)进行可认证的运行时监控。监控器需从图像中推断安全相关量,并提供有限样本保证,同时具备可复用性:一旦训练并校准,即可验证目标片段中的任意公式而无需逐公式重训。对于由有限词汇表生成的逻辑片段,我们证明了语义基(即各原子鲁棒性得分的向量)是单调、1-利普希茨可复用接口类中的最小预测目标:任何公式均可通过解析树导出的确定性解码器评估,且仅需一次置信校准即可覆盖整个片段,无需联合界。我们还引入了滚动预测监控器,仅预测当前谓词值并在线重构时序历史;该方法学习更易,但长期使用会趋于保守。在行人过街基准测试中,滚动监控在短时程下表现更优,而语义基监控在长时程下最紧致,可达4倍优势。我们在真实世界Waymo驾驶数据上验证了两种监控器,两者均在实证上满足置信覆盖保证。
原文摘要 · Abstract (English)
We study certified runtime monitoring of past-time signal temporal logic (ptSTL) from visual observations under partial observability. The monitor must infer safety-relevant quantities from images and provide finite-sample guarantees, while being \emph{reusable}: once trained and calibrated, it should certify any formula in a target fragment without per-formula retraining. For fragments induced by a finite dictionary of temporal atoms, we prove that the \emph{semantic basis}, the vector of atom robustness scores, is the minimum prediction target within the class of monotone, 1-Lipschitz reusable interfaces: any formula is evaluated by a deterministic decoder derived from the parse tree, and a single conformal calibration pass certifies the entire fragment with no union bound. We also introduce a \emph{rolling prediction monitor} that predicts only current predicate values and reconstructs temporal history online; this is easier to learn but grows conservative at long horizons. On a pedestrian-crossroad benchmark, rolling achieves tighter certified bounds at short horizons while the semantic-basis monitor is up to 4-times tighter at long horizons. We validate the presented monitors on real-world Waymo driving data, where both monitors satisfy the conformal coverage guarantee empirically.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。