arXiv:2608.16564cs.AI2026-08

基于情境的机器学习安全性能评估框架,提升系统可靠性。

CUBICS: Situation-aware performance estimation for safety-relevant ML components

论文配图:CUBICS: Situation-aware performance estimation for safety-relevant ML components
图 1 · 摘自论文原文
  • 按运行情境划分操作域,分组件建模性能风险
  • 利用主观逻辑动态更新各情境下的概率保障
  • 适合需要实测数据支持的安全关键系统验证

机器学习(ML)是当前创新的核心技术,但在安全相关应用中确保其安全性仍是一大挑战。一种有前景的方法是通过实地数据构建‘已验证使用’的论证,例如以影子模式或在安全边界内运行ML组件(MLCs),使其输出作为不影响安全的‘安全探针’被监控,进而以贝叶斯方式建立对现场性能的统计论证。然而,现有许多安全工程中的贝叶斯场数据方法将故障建模为单一全局失败概率的伯努利(或二项式)过程,且假设独立同分布,这通常不足以反映实际中性能高度依赖上下文的MLC。此外,统计证据还需涵盖相关情境(包括边缘情况),而构建整个系统的统一统计模型通常不可行。为此,本文提出CUBICS——一种面向组件、情境感知的高安全关键性ML组件性能估计框架。该框架将操作设计域划分为若干情境,并为每个安全相关组件定义一组情境特定的假设与概率保障,用主观逻辑(SL)表示并以贝叶斯方式更新。通过结合各情境的发生频率信念,CUBICS在无需整体系统级统计模型的前提下,推导出每个组件的总体风险估计,从而为基于场数据的模块化安全保证提供可复用的构建模块。

原文摘要 · Abstract (English)

Machine learning (ML) is a key technology driving innovation today, but ensuring ML safety remains a major challenge for safety-related applications. A promising idea is to build proven-in-use arguments from field data, e.g. by running ML components (MLCs) in shadow mode or within safety envelopes so that their outputs can be monitored as 'safe probes' without affecting safety. These probes can then be used to build a statistical argument about field performance in a Bayesian way. However, many Bayesian field-data approaches in safety engineering model failures as a simple Bernoulli (or binomial) process with a single global failure probability and i.i.d. trials, which is rarely adequate for MLCs whose performance depends strongly on context. Statistical evidence is also about coverage of relevant situations, including edge cases, and building a single integrated statistical model for the entire system is usually not feasible. To address these challenges, this paper introduces CUBICS, a context-modular framework for per-component, situation-aware performance estimation of safety-relevant ML components. CUBICS partitions the operational design domain into situations and, for each safety-relevant component, defines a set of situation-specific assumptions and probabilistic guarantees that are represented and updated in a Bayesian manner using Subjective Logic (SL). By combining these guarantees with beliefs about how often each situation occurs, CUBICS derives an overall risk estimate for each component without requiring a monolithic system-level statistical model, and thus provides a building block for modular, field-data based safety assurance.

机器学习安全情境感知贝叶斯推理可靠性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。