为安全研究制定新标准,防止灾难性失败。
Epistemic Norms for AI Safety and Alignment Research
- 以杜绝危险行为为目标,而非追求性能提升
- 识别出五大研究缺口,包括缺乏独立验证
- 提出可审计的伦理框架,兼顾透明与风险
主流AI研究注重能力增长,容忍低故障率;而安全与对齐研究的核心使命是确保在证据稀少、对抗环境和长尾风险下,灾难性失败永不发生。我们指出,二者分属两个独立维度:一是‘能力画像’——证明无危害行为的存在,而非具备正向能力;二是‘风险画像’——约束最坏情况下的结果,而非优化平均表现。现有主流认识论实践在这两方面均不充分。基于预注册的文献计量基线,我们识别出当前对齐研究中五个跨领域差距,包括制度化独立验证几乎缺失。为此,我们提出ECAISA(AI安全与对齐认识论规范),包含八项原则、三级评分体系、四级披露阶梯(平衡透明度与信息危害及商业机密)、分层适用方案、信息危害裁定流程及七项防操纵机制。ECAISA不认证系统安全,而是规范安全相关研究主张的记录、核查与依赖方式,以可审计性取代认证作为治理目标。
原文摘要 · Abstract (English)
Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment research has a different mission: to ensure that catastrophic failures never occur, under sparse evidence, adversarial dynamics, and fat-tailed risk. We argue that the two domains differ along two analytically independent axes---{\it capability profile}, demonstrating the absence of hazardous behaviours rather than the presence of positive capabilities, and {\it risk profile}, bounding worst-case outcomes under fat-tailed uncertainty rather than optimising average-case performance---and that mainstream epistemic practices are inadequate on both. Building on a structured synthesis grounded in a preregistered bibliometric baseline, we identify five cross-cutting gap dimensions in current alignment research, including the near-absence of institutionalised independent verification. To address these gaps, we propose {\sc ECAISA}, an Epistemic Code for AI Safety and Alignment comprising eight principles, a three-level scoring rubric, a four-level disclosure ladder that reconciles transparency with information-hazard and commercial-confidentiality constraints, a tiered applicability scheme, an information-hazard adjudication procedure, and seven anti-gaming mechanisms. {\sc ECAISA} does not certify that any AI system is safe; it constrains how safety-relevant research claims are documented, checked, and relied upon, with auditability rather than certification as its governance target.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。