arXiv:2605.25534cs.AI2026-05ACL

发现多模态大模型在复杂结构推理中会因认知过载导致安全失效

StructBreak: Structural Cognitive Overload-Induced Safety Failures in MLLMs

论文配图:StructBreak: Structural Cognitive Overload-Induced Safety Failures in MLLMs
图 1 · 摘自论文原文
  • 提出自动化框架StructBreak,量化结构认知过载现象
  • 92%平均攻击成功率,最高达97%,可黑盒触发有毒输出
  • 揭示安全机制在复杂推理下的失效,适合安全与可信AI研究者

多模态大语言模型(MLLM)虽擅长结构化推理,但存在严重的逻辑脆弱性,表现为结构认知过载(SCO),这是深度推理与安全对齐冲突的副产物。现有研究多关注像素级或文本扰动,忽视了SCO。为此,本文提出StructBreak——一个端到端自动化框架,用于量化SCO。利用该框架,我们发现一种新型高阶认知过载攻击,可在无需模型内部访问的黑盒场景下生效。基于此,构建覆盖十类威胁场景的综合基准。在六款主流MLLM上的实证评估显示,SCO极易引发毒性生成,平均攻击成功率达92%(Gemini 2.5最高达97%)。通过注意力动态、潜在空间拓扑与几何分析等多层次解释,发现StructBreak构成绕过安全过滤的新结构性通道。当前对齐范式在复杂多模态推理下效果有限,难以应对此类挑战。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) excel at structural reasoning yet suffer from a sharp logical brittleness in structural consistency. We term this phenomenon Structural Cognitive Overload (SCO), a byproduct of the contention between deep reasoning and safety alignment. However, prior work has predominantly targeted typographic and pixel-level perturbations, leaving the study of SCO largely unexplored. To this end, we propose StructBreak, an automated end-to-end framework designed to quantify SCO. By leveraging StructBreak, we uncover a novel higher-order cognitive overload attack paradigm; notably, this attack operates under a practical black-box setting, requiring no internal model access. Consequently, we utilize this framework to establish a comprehensive benchmark spanning ten diverse threat scenarios. Empirical evaluations on six leading MLLMs reveal that SCO readily triggers toxic generation, yielding a 92% average ASR (up to 97% on Gemini 2.5). To elucidate the mechanism of SCO, we further conduct model-level interpretations spanning attention dynamics, latent space topology, and geometric analysis. Our findings reveal that StructBreak acts as a novel structural channel to circumvent safety filters. Furthermore, the limited efficacy of inherent safety mechanisms underscores that current alignment paradigms are insufficient for the era of complex multimodal reasoning.

多模态模型安全漏洞认知过载黑盒攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。