构建大规模AI风险事件数据集,助力实证研究与安全治理
RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-dimensional annotations

- 从多语言新闻中自动提取并结构化AI风险事件
- 涵盖数亿条原始记录,形成可分析的事件集群与标注数据集
- 适合研究者开展AI安全、风险评估与跨源分析
随着人工智能系统在社会关键领域部署增多,与AI相关的危害和故障报告日益频繁且多样化。尽管现有治理框架提出了负责任AI的高层原则,但用于追踪和分析真实世界AI风险事件的大规模实证资源仍十分有限。现有事件集合多为人工整理,规模较小,难以支撑持续的数据驱动监测与下游计算分析。为此,我们提出RiskNet,一个基于大规模多语言新闻源构建的AI风险事件大数据集。RiskNet采用结构化流程实现AI风险新闻识别、事件级报告筛选、事件对齐与多维分类。该资源将分散的新闻报道整合为以事件为中心的记录,并提供事件分类、事件对齐与事件级风险标注的基准数据集。当前版本覆盖数亿条源记录,生成大规模AI风险相关报告集合,包含对齐的事件簇与标注子集。数据集亦可通过在线平台浏览与探索。本文详述数据来源、处理流程、分类体系设计及技术验证。RiskNet旨在支持AI安全、治理、风险分析与基准测试的下游研究,以及纵向与跨源的AI相关危害分析。通过提供结构化、可复用的实证资源,帮助弥合高层治理原则与实际风险事件之间的差距。
原文摘要 · Abstract (English)
As artificial intelligence (AI) systems are increasingly deployed across socially consequential domains, reports of AI-related harms and failures have grown in frequency and diversity. Although existing governance frameworks articulate high-level principles for responsible AI, large-scale empirical resources for tracking and analyzing real-world AI risk incidents remain limited. Existing incident collections are often manually curated, relatively small in scale, and insufficient for continuous, data-driven monitoring and downstream computational analysis. To address this need, we present RiskNet, a large-scale dataset of AI risk incidents constructed from large-scale multilingual news sources. RiskNet applies a structured pipeline for AI risk news identification, event-level report screening, incident alignment, and multi-dimensional incident classification. The resulting resource organizes dispersed news reports into incident-centered records and provides benchmark datasets for event classification, incident alignment, and incident-level risk labeling. In its current release, RiskNet covers hundreds of millions of source records and yields a large-scale collection of AI risk-related reports, including aligned incident clusters and annotated benchmark subsets. The dataset is also accessible through an online platform for browsing and exploration. We describe the data sources, processing workflow, taxonomy design, and technical validation of the resource. RiskNet is intended to support downstream research on AI safety, governance, risk analysis, and benchmarking, as well as longitudinal and cross-source analyses of AI-related harms. By providing a structured and reusable empirical resource, RiskNet helps bridge the gap between high-level governance principles and the documented realities of AI risk incidents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。