用AI检验AI,提升自动驾驶等关键系统的可靠性
Fighting AI with AI: Leveraging Foundation Models for Assuring AI-Enabled Safety-Critical Systems
- 用大语言模型将模糊需求转化为可验证的正式规范
- 用多模态模型以人类理解的概念测试视觉感知系统
- 适合从事AI安全验证的工程师和研究者
将深度神经网络(DNN)集成到航空航天、自动驾驶等安全关键系统中,带来保证方面的根本性挑战。AI系统的不可解释性,以及高层需求与底层网络表示之间的语义鸿沟,阻碍了传统验证方法的应用。这些AI特有挑战还受到需求工程长期问题的加剧,如自然语言规格的模糊性及形式化过程的可扩展性瓶颈。本文提出一种利用AI自身解决这些问题的方法,包含两个互补组件:REACT(基于AI的需求工程与一致性测试)利用大语言模型(LLMs),弥合非正式自然语言需求与形式化规范之间的差距,实现早期验证与确认;SemaLens(基于大模型的视觉感知语义分析)利用视觉语言模型(VLMs),使用人类可理解的概念对基于DNN的感知系统进行推理、测试与监控。二者共同构建了从非正式需求到可验证实现的完整流程。
原文摘要 · Abstract (English)
The integration of AI components, particularly Deep Neural Networks (DNNs), into safety-critical systems such as aerospace and autonomous vehicles presents fundamental challenges for assurance. The opacity of AI systems, combined with the semantic gap between high-level requirements and low-level network representations, creates barriers to traditional verification approaches. These AI-specific challenges are amplified by longstanding issues in Requirements Engineering, including ambiguity in natural language specifications and scalability bottlenecks in formalization. We propose an approach that leverages AI itself to address these challenges through two complementary components. REACT (Requirements Engineering with AI for Consistency and Testing) employs Large Language Models (LLMs) to bridge the gap between informal natural language requirements and formal specifications, enabling early verification and validation. SemaLens (Semantic Analysis of Visual Perception using large Multi-modal models) utilizes Vision Language Models (VLMs) to reason about, test, and monitor DNN-based perception systems using human-understandable concepts. Together, these components provide a comprehensive pipeline from informal requirements to validated implementations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。