首次系统绘制大模型安全失效区域地图,揭示不同模型的漏洞分布特征。
Manifold of Failure: Behavioral Attraction Basins in Language Models
- 用进化算法搜索模型行为偏差区域,定位多种脆弱性模式。
- 覆盖最高63%的失效行为,发现最多370个独特漏洞类型。
- 适合研究模型安全、对抗攻击与鲁棒性评估的学者参考。
现有研究多聚焦于将对抗样本映射回自然数据流形以恢复安全性,但我们认为全面理解AI安全需先刻画不安全区域本身。本文提出框架,系统化映射大语言模型(LLMs)的失效流形。将漏洞搜索重构为质量多样性问题,采用MAP-Elites算法揭示这些失效区域的连续拓扑结构,称为行为吸引盆地。使用对齐偏差(Alignment Deviation)作为质量度量,引导搜索至模型行为偏离预期对齐最严重的区域。在Llama-3-8B、GPT-OSS-20B和GPT-5-Mini三款模型上,MAP-Elites实现最高63%的行为覆盖率,发现最多370个不同的脆弱性生态位。结果揭示显著的模型特异性拓扑特征:Llama-3-8B呈现近乎普遍的漏洞平台(平均对齐偏差0.93),GPT-OSS-20B表现出碎片化结构且盆地空间聚集(平均0.73),而GPT-5-Mini展现出强鲁棒性,其对齐偏差上限仅为0.50。本方法生成可解释的全局安全景观图,是现有攻击方法(GCG、PAIR、TAP)无法提供的,推动范式从发现离散故障转向理解其内在结构。
原文摘要 · Abstract (English)
While prior work has focused on projecting adversarial examples back onto the manifold of natural data to restore safety, we argue that a comprehensive understanding of AI safety requires characterizing the unsafe regions themselves. This paper introduces a framework for systematically mapping the Manifold of Failure in Large Language Models (LLMs). We reframe the search for vulnerabilities as a quality diversity problem, using MAP-Elites to illuminate the continuous topology of these failure regions, which we term behavioral attraction basins. Our quality metric, Alignment Deviation, guides the search towards areas where the model's behavior diverges most from its intended alignment. Across three LLMs: Llama-3-8B, GPT-OSS-20B, and GPT-5-Mini, we show that MAP-Elites achieves up to 63% behavioral coverage, discovers up to 370 distinct vulnerability niches, and reveals dramatically different model-specific topological signatures: Llama-3-8B exhibits a near-universal vulnerability plateau (mean Alignment Deviation 0.93), GPT-OSS-20B shows a fragmented landscape with spatially concentrated basins (mean 0.73), and GPT-5-Mini demonstrates strong robustness with a ceiling at 0.50. Our approach produces interpretable, global maps of each model's safety landscape that no existing attack method (GCG, PAIR, or TAP) can provide, shifting the paradigm from finding discrete failures to understanding their underlying structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。