研究大模型安全过滤的极限:即使强化对齐与过滤,有害输出仍无法彻底消除。
On the Limits of Support-Preserving Alignment and Bounded Filtering
- 用支持集保持对齐+有限过滤,模拟真实安全机制
- 实测多模型在对抗提示下有害输出率始终高于零
- 揭示安全过滤存在不可逾越的理论与实证下限
我们研究了基于重塑基础模型输出分布的对齐方法,结合有限安全过滤,能否在现代大语言模型中将有害行为概率降至零。尽管近期研究指出偏好对齐下有害行为仍会持续,且外部过滤在最坏情况下计算困难,但尚不清楚在基本保留内部表征的实用对齐流程中,是否能完全消除有害行为而非仅压制其明显表现。本文通过支持集保持对齐算子与有限过滤算法,在黑盒、白盒和统计查询访问下形式化该场景,并分析其逼近理想消弭器的能力。基于此框架,我们提供计算与信息论论证,表明在这些约束下,有限过滤可能无法清除被基础模型分布支持的所有有害输出。为验证这些极限,我们在OpenRouter上对一系列先进开源与托管LLM进行评估,使用来自精选网络安全场景和PKU-SafeRLHF的对抗性提示,在有限黑盒、白盒及统计查询过滤下测试。结果显示,无论模型类型、过滤类别或查询预算如何,有害输出率随额外过滤计算下降但始终稳定在零以上,表明存在持续的实证危害下限。
原文摘要 · Abstract (English)
We study whether alignment schemes that reshape a base model's output distribution, combined with bounded safety filters, can drive the probability of harmful behavior to zero in modern large language models. Recent research suggests that harmful behaviors can persist under preference-based alignment and that external filtering can be computationally hard in the worst case, but it remains unclear whether practical alignment pipelines that largely preserve internal representations can eliminate harmful behavior entirely rather than merely suppressing its most visible forms. We formalize this setting using support-preserving alignment operators together with bounded filtering algorithms under black-box, white-box, and statistical-query access, and analyze their ability to approximate an ideal eliminator that removes all harmful mass. Building on this framework, we provide computational and information-theoretic arguments indicating that, under these constraints, bounded filtering may fail to eliminate all harmful outputs supported by the base model's distribution. To evaluate these limits empirically, we analyze a range of state-of-the-art open-weight and hosted LLMs accessed via OpenRouter under bounded black-box, white-box, and statistical-query filters on adversarial prompts drawn from curated cybersecurity scenarios and PKU-SafeRLHF. Across models, filter classes, and query budgets, the estimated harmful-output rate decreases with additional filtering compute but consistently plateaus above zero, suggesting a persistent empirical harm floor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。