构建可验证安全网,应对AI系统无声失效问题
Taming Silent Failures: A Framework for Verifiable AI Reliability
- 结合离线形式化合成与在线运行时监控,形成双重保障
- 在自动驾驶感知系统中检测出93.5%的隐性安全漏洞
- 符合ISO 26262标准,适合可靠性工程落地应用
人工智能(AI)融入高安全性系统带来了新的可靠性挑战:无声失效——即AI产生自信但错误的输出,可能造成危险。本文提出形式化保障与监测环境(FAME),将离线形式化合成的数学严谨性与在线运行时监测的实时警觉性相结合,为不透明的AI组件构建可验证的安全屏障。我们在自动驾驶感知系统中验证了其有效性,结果显示FAME成功检测出93.5%原本无法察觉的关键安全违规。通过将框架置于ISO 26262和ISO/PAS 8800标准背景下,我们为可靠性工程师提供了可认证、可部署的可信AI实现路径。FAME标志着从接受概率性能向确保可证明安全性的关键转变。
原文摘要 · Abstract (English)
The integration of Artificial Intelligence (AI) into safety-critical systems introduces a new reliability paradigm: silent failures, where AI produces confident but incorrect outputs that can be dangerous. This paper introduces the Formal Assurance and Monitoring Environment (FAME), a novel framework that confronts this challenge. FAME synergizes the mathematical rigor of offline formal synthesis with the vigilance of online runtime monitoring to create a verifiable safety net around opaque AI components. We demonstrate its efficacy in an autonomous vehicle perception system, where FAME successfully detected 93.5% of critical safety violations that were otherwise silent. By contextualizing our framework within the ISO 26262 and ISO/PAS 8800 standards, we provide reliability engineers with a practical, certifiable pathway for deploying trustworthy AI. FAME represents a crucial shift from accepting probabilistic performance to enforcing provable safety in next-generation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。