arXiv:2410.22526cs.AIcs.HC2024-10被引 11

用系统安全方法分析AI风险,发现组件交互带来的隐藏危害。

From Silos to Systems: Process-Oriented Hazard Analysis for AI Systems

  • 将系统安全框架STPA adapted为AI专用的PHASE方法
  • 在三种AI系统中识别出因组件交互引发的系统级风险
  • 适合关注AI治理与责任追溯的技术管理者

为有效应对人工智能系统可能带来的危害,必须识别并缓解系统级风险。现有分析方法仅孤立考察训练数据或模型等单个组件,忽视了组件间相互作用或其在企业开发流程中的位置所引发的风险。本文借鉴系统安全领域思想,将成熟的系统理论过程分析(STPA)框架应用于AI开发与运行过程。聚焦依赖机器学习算法的系统,对线性回归、强化学习及基于Transformer的生成模型三个案例开展STPA分析。研究探讨了控制与系统理论视角在AI系统中的适用性,并评估模型不可解释性、能力不确定性与输出复杂性等独特属性是否需对框架进行调整。结果表明,STPA的核心概念与步骤适用于AI系统,但需针对不同案例中的特定挑战进行定向适配。为此,本文提出面向AI系统的流程化危险分析(PHASE)指南。使用PHASE指南执行与解读STPA,可为负责管理AI风险的分析师提供四大关键支持:1)检测系统级风险,包括多个分散问题累积导致的隐患;2)明确承认社会因素对算法伤害的贡献;3)建立风险与可干预方之间的可追溯问责链条;4)持续监控并缓解新出现的风险。

原文摘要 · Abstract (English)

To effectively address potential harms from Artificial Intelligence (AI) systems, it is essential to identify and mitigate system-level hazards. Current analysis approaches focus on individual components of an AI system, like training data or models, in isolation, overlooking hazards from component interactions or how they are situated within a company's development process. To this end, we draw from the established field of system safety, which considers safety as an emergent property of the entire system. In this work, we translate System Theoretic Process Analysis (STPA) - a recognized system safety framework - for analyzing AI development and operation processes. We focus on systems that rely on machine learning algorithms and conduct STPA on three case studies involving linear regression, reinforcement learning, and transformer-based generative models. Our analysis explored how STPA's control and system-theoretic perspectives apply to AI systems and whether unique AI traits - such as model opacity, capability uncertainty, and output complexity - necessitate modifications to the framework. We find that the key concepts and steps of conducting an STPA apply to AI systems but require targeted adaptations to address AI-specific challenges that arise to differing degrees across three case studies. We present the Process-oriented Hazard Analysis for AI Systems (PHASE) as a guideline that adapts STPA concepts for AI. Applying and interpreting STPA using the PHASE guidelines enables four key affordances for analysts responsible for managing AI system harms: 1) detection of system-level hazards, including those from accumulation of disparate issues; 2) explicit acknowledgment of social factors contributing to algorithmic harms; 3) creation of traceable accountability chains between harms and those who can mitigate them; and 4) ongoing monitoring and mitigation of new hazards.

AI安全系统风险责任追溯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。