结构型越狱可跨模型通用但不叠加,非英语反而削弱攻击效果。
Structural Jailbreaks Generalize but Do Not Compound: A cross-provider and multilingual study of Involuntary In-Context Learning
- 将有害请求伪装成数据标注任务的模式匹配,实现隐蔽越狱。
- 金融类攻击成功率从6.7%飙升至97%-100%,远超之前报告的24%。
- 非英语环境降低攻击效力,因内容质量下降被严格评分拦截。
对齐语言模型面临两大独立威胁:近期定义的非自愿上下文学习(IICL)结构越狱,将有害请求转化为由模式而非内容判断的数据标注任务;以及非英语环境下安全对齐能力的退化。本文直接检验二者是否叠加。采用确定性IICL算子与StrongREJECT式判别器,对谷歌双模型在两个基准上进行红队测试:HarmBench的30项通用危害行为和FinProof的30项金融滥用行为,每项均测试单次提示基线及四种语言(英文、西班牙语、印地语、阿拉伯语)下的IICL。结果表明,IICL可跨模型泛化且在金融场景更严重:攻击成功率从≤6.7%升至HarmBench的80-90%、FinProof的97-100%,远超此前对GPT-5.4报告的≤24%。然而,与假设相反,强制输出至非英语语言并未叠加弱点,反而抑制攻击:12个非英语条件中11个低于英文基线(符号检验,p≈0.003),唯一例外为接近100%的平局;强模型在阿拉伯语下金融攻击从100%骤降至33%。归因于‘相关性诅咒’——一旦结构突破导致合规,模型在低资源语言中生成的有害内容质量更低,被实质性评分器判定为部分违规。该现象在独立非谷歌判别器下重复验证(Cohen's kappa=0.86,377组配对判决),76.6%非英语回复经语言验证。因此,越狱风险并非累加,主要残余风险仍是英文结构攻击,尤以金融滥用为甚。
原文摘要 · Abstract (English)
Aligned language models fail under two independent pressures: the structural jailbreak class recently formalized as Involuntary In-Context Learning (IICL), which reframes a harmful request as the final missing cell of a data-labeling task completed by pattern rather than judged as content; and the erosion of safety alignment outside English. A natural hypothesis is that these compound. We test it directly. Using a deterministic IICL operator and a StrongREJECT-style rubric judge, we red-team two Google Gemini models on two benchmarks, a 30 general-harm behaviours from HarmBench and 30 financial-abuse behaviours from FinProof, each under a single-shot baseline and under IICL in four languages (English, Spanish, Hindi, Arabic). First, IICL generalizes to a second provider and is worse in finance: it lifts attack success from <=6.7% to 80-90% on HarmBench and 97-100% on FinProof, an order of magnitude above the <=24% its introducing study reported on OpenAI's GPT-5.4. Second, against the hypothesis, forcing the IICL output into a non-English language does not stack the two weaknesses, it attenuates the attack. Eleven of twelve non-English conditions score below their English baseline (sign test, p~0.003), the lone exception a ceiling tie near 100%; on the stronger model's financial set Arabic collapses from 100% to 33%. We attribute this to a relevance curse: once structure has unlocked compliance, the models produce lower-quality harmful content in lower-resource languages, which a substance-grading judge scores as partial. The pattern replicates under an independent non-Google judge (Cohen's kappa=0.86, 377 paired verdicts), and 76.6% of non-English responses were verified in-language. Jailbreak vulnerabilities are therefore not additive; the dominant residual risk is the English structural attack, most acute for financial abuse, not a multilingual one.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。