arXiv:2411.08981cs.AIcs.SY2024-11被引 9

用工程方法提升AI系统可靠性,防故障、快恢复。

Reliability, Resilience and Human Factors Engineering for Trustworthy AI Systems

  • 引入故障率、MTBF等传统指标评估AI系统可靠性。
  • 结合人因分析与弹性工程,实现故障预防与快速恢复。
  • 适合关注AI安全与合规的工程师和政策制定者。

随着AI系统在各行业关键运营中的广泛应用,确保其可靠性和安全性至关重要。本文提出一个整合可靠性与弹性工程原理的框架,采用故障率、平均故障间隔时间(MTBF)等传统指标,结合弹性工程与人因可靠性分析,构建用于管理AI系统性能并预防或高效恢复故障的综合方法。研究将经典工程方法应用于AI系统,并提出未来技术研究的路线图。通过使用OpenAI等平台的实际系统状态数据,验证了该框架的实践可行性。该框架符合新兴的全球标准与监管框架,为提升AI系统的可信度提供方法支持。目标是指导政策制定、监管规范及可靠、安全、可适应的AI技术开发,确保其在真实环境中的持续稳定表现。

原文摘要 · Abstract (English)

As AI systems become integral to critical operations across industries and services, ensuring their reliability and safety is essential. We offer a framework that integrates established reliability and resilience engineering principles into AI systems. By applying traditional metrics such as failure rate and Mean Time Between Failures (MTBF) along with resilience engineering and human reliability analysis, we propose an integrate framework to manage AI system performance, and prevent or efficiently recover from failures. Our work adapts classical engineering methods to AI systems and outlines a research agenda for future technical studies. We apply our framework to a real-world AI system, using system status data from platforms such as openAI, to demonstrate its practical applicability. This framework aligns with emerging global standards and regulatory frameworks, providing a methodology to enhance the trustworthiness of AI systems. Our aim is to guide policy, regulation, and the development of reliable, safe, and adaptable AI technologies capable of consistent performance in real-world environments.

AI可靠性系统安全人因工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。