证明了大模型安全过滤无法仅靠外部机制实现,智能与判断不可分割。
On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment
- 提出对抗性提示在计算上无法与正常提示区分,导致输入过滤失效
- 证明在特定条件下输出过滤在计算上不可行,需依赖模型内部机制
- 强调黑箱访问无法保障安全,需将判断嵌入智能核心
随着大语言模型(LLMs)的广泛应用,其生成有害内容的潜在风险引发关注。本文研究对齐挑战,重点分析用于阻止不安全信息生成的过滤机制。我们发现输入提示和输出内容的过滤均存在计算障碍:首先,存在某些LLM,其对抗性提示可轻易构造,且在计算上与良性提示无法区分,导致高效提示过滤器不存在;其次,在自然设定下,输出过滤被证明是计算上不可行的。所有分离结果基于密码学难假设。此外,我们形式化并研究了放松的缓解方法,进一步揭示了计算限制。结论表明,安全无法通过外部设计的过滤器实现;特别是,仅具备对模型的黑盒访问权限不足以保证安全。基于技术结果,我们认为对齐人工智能系统的智能无法与其判断分离。
原文摘要 · Abstract (English)
With the increased deployment of large language models (LLMs), one concern is their potential misuse for generating harmful content. Our work studies the alignment challenge, with a focus on filters to prevent the generation of unsafe information. Two natural points of intervention are the filtering of the input prompt before it reaches the model, and filtering the output after generation. Our main results demonstrate computational challenges in filtering both prompts and outputs. First, we show that there exist LLMs for which there are no efficient prompt filters: adversarial prompts that elicit harmful behavior can be easily constructed, which are computationally indistinguishable from benign prompts for any efficient filter. Our second main result identifies a natural setting in which output filtering is computationally intractable. All of our separation results are under cryptographic hardness assumptions. In addition to these core findings, we also formalize and study relaxed mitigation approaches, demonstrating further computational barriers. We conclude that safety cannot be achieved by designing filters external to the LLM internals (architecture and weights); in particular, black-box access to the LLM will not suffice. Based on our technical results, we argue that an aligned AI system's intelligence cannot be separated from its judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。