为大模型内部使用风险提供统一报告框架,助力企业合规与安全管控。
Risk Reporting for Developers' Internal AI Model Use
- 基于自主行为与内部威胁两类风险,拆解手段、动机、机会三要素。
- 要求每轮高风险模型内测后提交报告,明确安全理由与残留风险。
- 适配美欧多国法规,供安全团队、监管方参考执行。
前沿AI公司通常在公开发布前数周至数月内部测试其最先进模型,例如Anthropic的Mythos Preview模型在内部使用至少六周。此类内部使用带来外部监管框架难以覆盖的风险。加州《前沿人工智能透明度法案》(SB 53)、纽约《负责任AI安全与教育法案》(RAISE)及欧盟通用人工智能行为准则均要求开发者制定并实施内部使用风险管理计划,并提交内部使用风险报告,说明防护措施与剩余风险。本指南提供一套协调一致的报告标准,适用于上述三项法规。主要面向前沿AI开发企业的评估与安全部门,次要对象为监管与审计人员。鉴于AI研发自动化加速且外部难以观测企业内部模型使用情况,定期详尽的风险报告可能是唯一可预见并管理潜在风险的机制。每当部署更强大或更高风险的模型时,开发者应生成风险报告,论证其安全性。报告框架围绕两大威胁向量——自主AI误行为与内部威胁,各设手段、动机、机会三个风险因子。
原文摘要 · Abstract (English)
Frontier AI companies first deploy their most advanced models internally, for weeks or months of safety testing, evaluation, and iteration, before a possible public release. For example, Anthropic recently developed a new class of model with advanced cyberoffense-relevant capabilities, Mythos Preview, which was available internally for at least six weeks before it was publicly announced. This internal use creates risks that external deployment frameworks may fail to address. Legal frameworks, notably California's Transparency in Frontier Artificial Intelligence Act (SB 53), New York's Responsible AI Safety And Education (RAISE) Act, and the EU's General-Purpose AI Code of Practice, all discuss risks from internal AI use. They require frontier developers to make and implement plans for how to manage risks from internal use, and to produce internal use risk reports describing their safeguards and any residual risks. This guide provides a harmonized standard for companies to produce internal use risk reports suitable for all three regulatory frameworks. It is addressed primarily to evaluation and safety teams at frontier AI developers, and secondarily to regulators and auditors seeking to understand what good reporting looks like. Given the pace of AI R&D automation and the limited external visibility into how companies use their most capable models internally, regular and detailed risk reporting may be one of the few mechanisms available to ensure that the risks from internal AI use are identified and managed before they materialize. Whenever a substantially more capable or riskier model is deployed internally, the developer should create a risk report and argue why the model is safe to deploy. We structure the reporting framework around two threat vectors -- autonomous AI misbehavior and insider threats -- and three risk factors for each: means, motive, and opportunity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。