用层级结构识别AI行为风险,提前发现潜在危险推理模式。
PRISM Risk Signal Framework: Hierarchy-Based Red Lines for AI Behavioral Risk
- 基于价值、证据、来源三层结构定义27种风险信号
- 通过双阈值机制区分确认风险与预警信号,准确率超90%
- 适合用于AI安全评估与模型可靠性审查
当前AI安全方法在具体案例层面设定红灯:特定提示、输出或危害。本文提出可在更根本的层面——支配AI推理的价值、证据与信息源层级——设置红灯。利用PRISM(基于画像的推理完整性堆栈测量)框架,我们构建了27种行为风险信号的分类体系,这些信号源于AI系统在价值优先级(L4)、证据权重(L3)和信息源信任度(L2)方面的结构性异常。每种信号通过绝对排名位置与相对胜率差距的双阈值原则评估,实现确认风险与预警信号的两级分类。该层级方法相较传统案例式红灯具备三大优势:前瞻性(在有害输出产生前识别危险推理结构)、全面性(单一价值层级信号涵盖无限案例违规)和可度量性(基于约39.7万次强制选择数据)。实验使用7个AI模型在三个权威层级上的近39.7万次强制选择响应,验证了该框架对具有极端结构特征、情境依赖风险及平衡层级模型的有效区分能力。
原文摘要 · Abstract (English)
Current approaches to AI safety define red lines at the case level: specific prompts, specific outputs, specific harms. This paper argues that red lines can be set more fundamentally -- at the level of value, evidence, and source hierarchies that govern AI reasoning. Using the PRISM (Profile-based Reasoning Integrity Stack Measurement) framework, we define a taxonomy of 27 behavioral risk signals derived from structural anomalies in how AI systems prioritize values (L4), weight evidence types (L3), and trust information sources (L2). Each signal is evaluated through a dual-threshold principle combining absolute rank position and relative win-rate gap, producing a two-tier classification (Confirmed Risk vs. Watch Signal). The hierarchy-based approach offers three advantages over case-specific red lines: it is anticipatory rather than reactive (detecting dangerous reasoning structures before they produce harmful outputs), comprehensive rather than enumerative (a single value-hierarchy signal subsumes an unlimited number of case-specific violations), and measurable rather than subjective (grounded in empirical forced-choice data). We demonstrate the framework's detection capacity using approximately 397,000 forced-choice responses from 7 AI models across three Authority Stack layers, showing that the signal taxonomy successfully discriminates between models with structurally extreme profiles, models with context-dependent risk, and models with balanced hierarchies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。