为航空避撞AI系统设计了一套评估数据代表性的方法。
Coverage-Driven Verification for Safety-by-Design in AI-Based Collision Avoidance Systems

- 从目标分布建模出发,构建可量化的覆盖率评估流程。
- 用KL散度和Cramér's V替代卡方检验,适配大规模数据场景。
- 适用于遵循EASA安全标准的AI系统开发与验证,尤其适合航空领域。
人工智能在航空系统中潜力巨大,但其在安全关键应用中的集成需满足严格的行业安全标准。针对基于AI与机器学习(ML)的系统,欧洲航空安全局(EASA)强调必须证明运行设计域(ODD)及其训练数据分布的代表性与完整性。然而,目前尚缺乏结构化工程流程来定义目标分布并评估其代表性。本文提出一种面向航空安全保证的AI/ML系统中各子模块ODD代表性的评估方法。该方法从系统性识别合适的目标分布开始,建立从ODD定义、参数分布建模到覆盖率定量评估与解释的全流程框架,符合EASA的学习保障目标。针对大规模数据场景,研究发现卡方拟合优度检验不适用,转而采用Kullback--Leibler散度与Cramér's $V$进行代表性评估。以基于AI的机载避撞系统为例,使用先前水平避撞系统(HCAS)与垂直避撞系统(VCAS)的仿真数据进行验证。结果表明,统计分布比较方法能有效支持对安全关键型AI应用的代表性评估,推动符合新兴EASA指南的系统性安全设计工程流程。
原文摘要 · Abstract (English)
Artificial Intelligence (AI) offers significant potential for future aviation systems; however, its integration into safety-critical applications requires compliance with the aviation sector's stringent safety standards. For AI and Machine Learning (ML)-based systems, the European Union Aviation Safety Agency (EASA) emphasizes the need to demonstrate the representativeness and completeness of the Operational Design Domain (ODD) and the associated data distributions used during development and verification. Despite this requirement, a structured engineering process for defining target distributions and evaluating representativeness within ODDs remains largely unexplored. This work presents a method for representativeness assessment of AI/ML constituent ODDs in the context of aviation safety assurance. Starting from the methodical identification of suitable target distributions, a process flow is proposed that guides developers from ODD definition and parameter distribution modeling to the quantitative assessment and interpretation of coverage results with respect to EASA's learning assurance objectives. As quantitative measures, the chi-squared goodness-of-fit test is examined and found unsuitable for the large data sets arising in this setting, leading to the adoption of the Kullback--Leibler divergence and Cramér's $V$ for the representativeness assessment. The method is demonstrated using the example of AI-based airborne collision avoidance, employing experimental data from previous Horizontal Collision Avoidance System (HCAS) and Vertical Collision Avoidance System (VCAS) simulations. The results illustrate how statistical distribution comparison methods can support the assessment of representativeness for safety-critical AI applications and contribute toward a systematic Safety-by-Design AI engineering process aligned with emerging EASA guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。