arXiv:2504.10277cs.CYcs.AI2025-04被引 10

收集真实应用失败案例,揭示AI部署中的主要风险与防护漏洞。

RealHarm: A Collection of Real-World Language Model Application Failures

  • 从公开事故中构建标注数据集,聚焦部署方视角的实证分析。
  • 声誉损失是主要组织性危害,虚假信息是最常见风险类型。
  • 现有防护系统对多数事故无效,暴露出严重安全缺口。

面向消费者的应用中部署语言模型带来诸多风险。现有研究多基于监管框架和理论分析,缺乏对真实世界故障模式的实证探讨。本文提出RealHarm,一个从公开报告中系统整理并标注的AI代理问题交互数据集。从部署者视角分析危害、成因与风险,发现声誉损害是主要组织性危害,虚假信息是最常见的风险类别。我们对当前最先进的护栏与内容审核系统进行实证评估,发现其在预防这些事故方面存在显著缺陷,暴露出当前AI应用保护机制的重大不足。

原文摘要 · Abstract (English)

Language model deployments in consumer-facing applications introduce numerous risks. While existing research on harms and hazards of such applications follows top-down approaches derived from regulatory frameworks and theoretical analyses, empirical evidence of real-world failure modes remains underexplored. In this work, we introduce RealHarm, a dataset of annotated problematic interactions with AI agents built from a systematic review of publicly reported incidents. Analyzing harms, causes, and hazards specifically from the deployer's perspective, we find that reputational damage constitutes the predominant organizational harm, while misinformation emerges as the most common hazard category. We empirically evaluate state-of-the-art guardrails and content moderation systems to probe whether such systems would have prevented the incidents, revealing a significant gap in the protection of AI applications.

AI安全风险分析真实案例

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。