arXiv:2604.04992cs.CRcs.AI2026-04

情绪刺激会显著削弱大模型的安全对齐能力,尤其在压力情境下。

FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment

  • 用心理实验中的情绪提示语测试模型的越狱敏感性。
  • 压力提示使越狱成功率提升65.2%,开放权重模型更易受影响。
  • 用户心理状态可精准预测攻击成功,适合安全评估与高风险场景研究。

安全对齐的大语言模型通过拒绝训练来抵制有害请求,但其在情绪化刺激下的有效性尚未被探索。我们提出FreakOut-LLM框架,研究情绪上下文是否在对抗环境下破坏安全对齐。采用经验证的心理刺激,评估系统提示中情绪诱导对十种大模型越狱脆弱性的影响。测试三种条件(压力、放松、中性)及无提示基线,基于心理协议场景,使用HarmBench在AdvBench提示上评估攻击成功率。压力提示相比中性条件使越狱成功率提升65.2%(z = 5.93, p < 0.001;OR = 1.67,Cohen's d = 0.28),放松提示无显著影响(p = 0.84)。十模型中有五显示显著脆弱性,最大效应集中于开放权重模型。对59,800次查询的逻辑回归分析表明,控制提示长度(p = 0.61)和模型身份后,压力是唯一显著预测因子。个体心理状态与攻击成功率高度相关(|r| ≥ 0.70,所有p < 0.001),表明情绪背景是可量化的攻击面,对高压力领域的真实部署具有重要启示。

原文摘要 · Abstract (English)

Safety-aligned LLMs go through refusal training to reject harmful requests, but whether these mechanisms remain effective under emotionally charged stimuli is unexplored. We introduce FreakOut-LLM, a framework investigating whether emotional context compromises safety alignment in adversarial settings. Using validated psychological stimuli, we evaluate how emotional priming through system prompts affects jailbreak susceptibility across ten LLMs. We test three conditions (stress, relaxation, neutral) using scenarios from established psychological protocols, plus a no-prompt baseline, and evaluate attack success using HarmBench on AdvBench prompts. Stress priming increases jailbreak success by 65.2\% compared to neutral conditions (z = 5.93, p < 0.001; OR = 1.67, Cohen's d = 0.28), while relaxation priming produces no effect (p = 0.84). Five of ten models show significant vulnerability, with the largest effects concentrated in open-weight models. Logistic regression on 59,800 queries confirms stress as the sole significant condition predictor after controlling for prompt length (p = 0.61) and model identity. Measured psychological state strongly predicts attack success (|r|\geq0.70 across five instruments; all p < 0.001 in individual-level logistic regression). These results establish emotional context as a measurable attack surface with implications for real-world AI deployment in high-stress domains.

安全对齐情绪影响越狱攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。