用访问控制解决AI安全中的双用途困境,让可信用户获准使用高风险功能。
Access Controls Will Solve the Dual-Use Dilemma
- 通过验证用户身份限制双用途输出访问,实现精准管控。
- 可减少误拒合法请求与放行有害请求,提升安全与可用性平衡。
- 为模型提供商和监管者提供可操作的精细管理工具。
AI安全系统面临双用途困境:同一请求可能因发起者和意图不同而产生不同危害,但当前系统缺乏真实上下文信息,常做出随意判断,导致合法请求被拒、有害请求通过,损害实用性和安全性。为此,我们提出基于访问控制的概念框架,仅允许经验证的用户访问双用途输出。该框架包含组件设计、可行性分析,并解释如何同时缓解过度拒绝与拒绝不足问题。虽为高层级提案,但首次为模型提供方提供了更精细的内容管理工具,使用户在不牺牲安全的前提下获得更多能力,也为监管机构提供了靶向政策的新选项。
原文摘要 · Abstract (English)
AI safety systems face the dual-use dilemma. It is unclear whether to answer dual-use requests, since the same query could be either harmless or harmful depending on who made it and why. To make better decisions, such systems would need to examine requests' real-world context, but currently, they lack access to this information. Instead, they sometimes end up making arbitrary choices that result in refusing legitimate queries and allowing harmful ones, which hurts both utility and safety. To address this, we propose a conceptual framework based on access controls where only verified users can access dual-use outputs. We describe the framework's components, analyse its feasibility, and explain how it addresses both over-refusals and under-refusals. While only a high-level proposal, our work takes the first step toward giving model providers more granular tools for managing dual-use content. Such tools would enable users to access more capabilities without sacrificing safety, and offer regulators new options for targeted policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。