现有安全机制无法防范模仿攻击,模型安全面临能力、可靠性和开放性的三难困境。
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs

- 区分模型能力与使用证据,识别可复制证据下的最大威胁
- 证明在可复制证据下,攻击者协助至少存在固定下限
- 提出可信凭证可突破限制,适合安全敏感场景的系统设计
大型语言模型的安全机制在看到答案实际用途前就决定是否回应,这在双重用途任务中带来根本问题:同一回答可能帮助合法用户或攻击者,而攻击者可模仿良性请求与交互历史。本文将模型释放的能力与下游使用证据分离,当该证据可复制时,推导出攻击者获得帮助的精确最坏情况下界,同时保留有用回答。结果揭示一个安全三难困境:有用能力、可靠安全与开放访问三者不可共存。接着表明,可信凭证可通过引入难以复制的、预测真实下游使用的信 息,补充现有防护机制,并确定消除下界的更强条件。双重用途评估、自适应攻击实验及已部署的受信任访问项目的数据支持这些条件的实际意义。
原文摘要 · Abstract (English)
Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and interaction history. We separate the capability released by the model from the evidence available about downstream use. When that evidence is copyable, we derive the exact worst-case floor on attacker assistance while preserving useful answers. The result yields a safety trilemma: Useful Capability, Reliable Safety, and Open Access cannot coexist. We then show how a trusted credential can complement existing safeguards by adding hard-to-copy information that predicts actual downstream use, and identify the stronger condition needed to eliminate the floor. Evidence from dual-use evaluations, adaptive attacks, and deployed trusted-access programs supports the practical relevance of these conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。