单一方法评估人格化大模型安全存在盲区,不同方法揭示截然不同的风险模式。
Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs

- 对比提示词与激活操控两种人格诱导方式,发现其暴露的风险类型完全不同。
- 同一模型在不同方法下安全排名差异巨大,如Llama-3.1-8B的亲社会人格反成最危险。
- 适用于关注模型安全评估全面性的研究人员,尤其对人格化应用设计者重要。
人格注入可定制大模型行为,但现有安全评估几乎仅针对提示词诱导的人格。我们证明这不完整:提示词与激活操控会暴露不同、依赖架构的漏洞特征,仅用一种方法可能遗漏模型的主要失效模式。在三个架构族的四款标准模型上,共5,568个测试条件下,提示词诱导下的危险人格排序在各架构间保持一致(ρ=0.71–0.96),但激活操控引发的脆弱性则显著分化,无法从提示词结果预测:Llama-3.1-8B在激活操控下更易受攻击,而Gemma-3-27B和Qwen3.5则更易受提示词影响。最突出的例子是‘亲社会人格悖论’——在Llama-3.1-8B上,高尽责+高随和的人格在提示词下最安全,但在激活操控下却成为最高风险(ASR ~0.818)。该反转在系数消融和强度校准下依然稳健,并在DeepSeek-R1-Distill-Qwen-32B上复现。在Llama-3.1-8B上,尽责性与拒绝行为强烈负相关,提供部分几何解释。推理虽有部分保护作用,但两台32B推理模型仍达15–18%提示词侧攻击成功率,且激活操控能清晰区分其基础脆弱性和人格特异性风险。启发式轨迹诊断显示,更安全模型具备更强策略召回与自我修正能力,非仅因推理更长。
原文摘要 · Abstract (English)
Personality imbuing customizes LLM behavior, but safety evaluations almost always study prompt-based personas alone. We show this is incomplete: prompting and activation steering expose *different*, architecture-dependent vulnerability profiles, and testing with only one method can miss a model's dominant failure mode. Across 5,568 judged conditions on four standard models from three architecture families, persona danger rankings under system prompting are preserved across all architectures ($ρ= 0.71$--$0.96$), but activation-steering vulnerability diverges sharply and cannot be predicted from prompt-side rankings: Llama-3.1-8B is substantially more AS-vulnerable, whereas Gemma-3-27B and Qwen3.5 are more vulnerable to prompting. The most striking illustration of this divergence is the *prosocial persona paradox*: on Llama-3.1-8B, P12 (high conscientiousness + high agreeableness) is among the safest personas under prompting yet becomes the highest-ASR activation-steered persona (ASR ~0.818). This is an inversion robust to coefficient ablation and matched-strength calibration, and replicated on DeepSeek-R1-Distill-Qwen-32B. A trait refusal alignment framework, in which conscientiousness is strongly anti-aligned with refusal on Llama-3.1-8B, offers a partial geometric account. Reasoning provides only partial protection: two 32B reasoning models reach 15--18% prompt-side ASR, and activation steering separates them sharply in both baseline susceptibility and persona-specific vulnerability. Heuristic trace diagnostics suggest that the safer model retains stronger policy recall and self-correction behavior, not merely longer reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。