arXiv:2504.00091cs.CYcs.AI2025-04被引 1

从信息本质出发,构建生成式AI风险评估框架。

A First-Principles Based Risk Assessment Framework and the IEEE P3396 Standard

  • 按信息层级分类生成内容,区分感知、知识、决策与控制四类风险
  • 强调结果风险优先,明确开发者、部署者等各方责任归属
  • 为监管标准提供理论基础,适合政策制定与安全设计参考

生成式人工智能在内容生成与决策支持中实现前所未有的自动化,但也带来新风险。本文提出基于第一性原理的风险评估框架,支撑IEEE P3396《AI风险、安全、可信与责任推荐实践》。区分过程风险(系统构建与运行中的风险)与结果风险(输出及其现实影响),主张治理应优先关注结果风险。核心是信息导向的本体论,将生成内容分为四类:感知级信息、知识级信息、决策/行动方案信息、控制令牌(访问或资源指令)。该分类使危害识别系统化,并依据生成信息类型精准划分开发、部署、使用、监管方的责任。每类信息对应不同后果风险(如欺骗、虚假信息、不安全建议、安全漏洞),需匹配特定风险指标与缓解措施。框架以信息本质、人类能动性与认知为基础,确保风险评估贴合生成内容对人类理解与行为的影响。相比泛化的应用分类,该方法更利于责任清晰与精准防护。文中附表展示信息类型与风险、责任的映射关系。本研究旨在为IEEE P3396及更广泛的AI治理提供严谨的第一性原理基础,支持负责任创新。

原文摘要 · Abstract (English)

Generative Artificial Intelligence (AI) is enabling unprecedented automation in content creation and decision support, but it also raises novel risks. This paper presents a first-principles risk assessment framework underlying the IEEE P3396 Recommended Practice for AI Risk, Safety, Trustworthiness, and Responsibility. We distinguish between process risks (risks arising from how AI systems are built or operated) and outcome risks (risks manifest in the AI system's outputs and their real-world effects), arguing that generative AI governance should prioritize outcome risks. Central to our approach is an information-centric ontology that classifies AI-generated outputs into four fundamental categories: (1) Perception-level information, (2) Knowledge-level information, (3) Decision/Action plan information, and (4) Control tokens (access or resource directives). This classification allows systematic identification of harms and more precise attribution of responsibility to stakeholders (developers, deployers, users, regulators) based on the nature of the information produced. We illustrate how each information type entails distinct outcome risks (e.g. deception, misinformation, unsafe recommendations, security breaches) and requires tailored risk metrics and mitigations. By grounding the framework in the essence of information, human agency, and cognition, we align risk evaluation with how AI outputs influence human understanding and action. The result is a principled approach to AI risk that supports clear accountability and targeted safeguards, in contrast to broad application-based risk categorizations. We include example tables mapping information types to risks and responsibilities. This work aims to inform the IEEE P3396 Recommended Practice and broader AI governance with a rigorous, first-principles foundation for assessing generative AI risks while enabling responsible innovation.

AI治理风险评估生成式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。