提出以人为本、按危害严重度自适应的AI风险评估框架。
AI Harmonics: a human-centric and harms severity-adaptive AI risk assessment framework
- 基于实证事故数据构建新型危害评估指标AIH,用等级数据衡量影响。
- 政治与物理危害最集中,前者削弱信任,后者威胁生命安全。
- 可识别危害分布不均,助力政策制定者精准施策。
人工智能的绝对主导带来了前所未有的社会危害与风险。现有评估模型多关注内部合规性,忽视多元利益相关方视角与真实世界后果。本文提出一种以人为本、危害严重度自适应的新范式,基于实证事故数据构建AI Harmonics框架。该框架包含新型人工智能危害评估指标(AIH),利用有序严重程度数据捕捉相对影响,无需精确数值估算。AI Harmonics融合稳健通用方法与数据驱动、利益相关方感知的框架,用于探索和优先排序AI危害。在标注事故数据上的实验表明,政治与物理危害最为集中,前者侵蚀公众信任,后者带来严重甚至致命风险,凸显本方法的现实意义。最终证明,该框架能持续识别危害分布不均现象,使政策制定者与组织可有效聚焦缓解措施。
原文摘要 · Abstract (English)
The absolute dominance of Artificial Intelligence (AI) introduces unprecedented societal harms and risks. Existing AI risk assessment models focus on internal compliance, often neglecting diverse stakeholder perspectives and real-world consequences. We propose a paradigm shift to a human-centric, harm-severity adaptive approach grounded in empirical incident data. We present AI Harmonics, which includes a novel AI harm assessment metric (AIH) that leverages ordinal severity data to capture relative impact without requiring precise numerical estimates. AI Harmonics combines a robust, generalized methodology with a data-driven, stakeholder-aware framework for exploring and prioritizing AI harms. Experiments on annotated incident data confirm that political and physical harms exhibit the highest concentration and thus warrant urgent mitigation: political harms erode public trust, while physical harms pose serious, even life-threatening risks, underscoring the real-world relevance of our approach. Finally, we demonstrate that AI Harmonics consistently identifies uneven harm distributions, enabling policymakers and organizations to target their mitigation efforts effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。