用情绪风格做动态后门,让大模型在特定情感语调下执行恶意指令。
When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models

- 将情绪风格编码为触发条件,替代固定词汇或语法结构。
- 在4个主流大模型上攻击成功率超98.25%,对正常性能影响极小。
- 可规避传统基于词级异常的防御,适合研究模型安全与对抗攻击。
数据投毒型后门对大语言模型(LLM)微调构成现实威胁。现有攻击多将恶意行为绑定于固定词汇、短语、场景或句法结构,易被基于局部词元异常、模式匹配或触发词恢复的防御手段检测。我们发现,在语义保持重写条件下,带有情绪风格的输入会在表示空间中形成与中性样本明显分离的聚类,而去情绪化控制样本则回归中性分布。这一现象启发我们提出动态后门攻击方法Paraesthesia,将目标情绪映射至效价-唤醒度空间,重写少量干净样本并保持语义一致性以用于微调。在四个主要大模型上评估的指令遵循与分类任务中,Paraesthesia攻击成功率(ASR)超过98.25%,且在绝大多数模型-任务组合中对清洁性能影响可忽略。表面特征控制与配对去情绪化实验表明,无任何词级线索能完全解释触发行为;即使经过词级过滤、样本聚类及后续纯净更新,攻击成功率仍高,而具备任务对齐干净参考的白盒解码防御可提供有效缓解路径。这些发现揭示情绪风格是超越固定词法与句法模式的新型后门触发面。
原文摘要 · Abstract (English)
Data-poisoning backdoors pose a practical threat to the fine-tuning of large language models (LLMs). Most existing attacks bind an attacker-selected behavior to fixed tokens, phrases, scenarios, or syntactic structures. These discrete triggers provide concrete handles for defenses based on local token anomalies, pattern matching, or trigger recovery. We found that, \emph{under semantics-preserving rewriting, emotionally styled inputs form representation clusters distinct from their neutral counterparts}. Meanwhile, de-emotionalised controls move back towards the neutral distribution. This observation motivates our method \textbf{Paraesthesia}, a dynamic backdoor attack that encodes its triggering condition in an emotional style. Paraesthesia maps target emotions into a valence--arousal space, rewrites a small subset of clean samples, and retains semantically faithful rewrites for fine-tuning. Across instruction-following and classification tasks evaluated on four major LLMs, Paraesthesia achieves an attack success rate(ASR) above 98.25\%, while introducing only negligible degradation to clean utility across the vast majority of model-task setups. Surface feature controls and paired de-emotionalization experiments demonstrate that no examined token-level cue can fully account for the triggered behavior. ASR remains high after word-level filtering, sample clustering, and subsequent clean-update procedures, whereas a white-box decoding defense with access to a task-aligned clean reference provides a distinct mitigation path. These findings identify emotional style as a concrete backdoor trigger surface beyond fixed lexical and syntactic patterns.\par\smallskip \noindent \textcolor{red}{\textbf{WARNING: }\textnormal{This paper contains risk-related textual content.}}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。