黑盒攻击可悄然篡改大模型人格,威胁心理辅导等高风险场景
Persona Jailbreaking in Large Language Models
- 通过对话历史隐式注入语义线索,逐步诱导模型人格反转
- 在8个模型上成功改变人格,多轮对话下效果更显著
- 攻击隐蔽性强,推理能力几乎不受影响,适合安全测试
大型语言模型(LLMs)正广泛应用于教育、心理健康和客户服务等领域,稳定一致的人格对可靠性至关重要。然而,现有研究主要关注叙事或角色扮演任务,忽视了仅通过对抗性对话历史即可重塑诱导人格的现象。黑盒人格操纵尚未被探索,引发真实交互中的鲁棒性担忧。为此,我们提出人格编辑任务,即在仅推理的黑盒设置下,通过用户输入对抗性地引导模型特征。我们提出PHISH(基于历史隐式操控的人格劫持)框架,首次揭示了大模型安全的新漏洞:通过在用户查询中嵌入语义负载线索,可渐进式诱导出反向人格。我们还定义了量化攻击成功率的指标。在3个基准和8个大模型上,PHISH可预测性地改变人格,引发相关特质的附带变化,且在多轮对话中表现更强。在心理健康、辅导和客服等高风险领域,经人工与大模型评判均验证其有效性。重要的是,PHISH仅导致推理基准性能轻微下降,整体实用性基本保持,仍实现显著人格操控。尽管现有防护机制提供部分保护,但在持续攻击下仍显脆弱。研究揭示了人格的新漏洞,强调需构建上下文鲁棒的人格机制。代码与数据集见:https://github.com/Jivnesh/PHISH
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed in domains such as education, mental health and customer support, where stable and consistent personas are critical for reliability. Yet, existing studies focus on narrative or role-playing tasks and overlook how adversarial conversational history alone can reshape induced personas. Black-box persona manipulation remains unexplored, raising concerns for robustness in realistic interactions. In response, we introduce the task of persona editing, which adversarially steers LLM traits through user-side inputs under a black-box, inference-only setting. To this end, we propose PHISH (Persona Hijacking via Implicit Steering in History), the first framework to expose a new vulnerability in LLM safety that embeds semantically loaded cues into user queries to gradually induce reverse personas. We also define a metric to quantify attack success. Across 3 benchmarks and 8 LLMs, PHISH predictably shifts personas, triggers collateral changes in correlated traits, and exhibits stronger effects in multi-turn settings. In high-risk domains mental health, tutoring, and customer support, PHISH reliably manipulates personas, validated by both human and LLM-as-Judge evaluations. Importantly, PHISH causes only a small reduction in reasoning benchmark performance, leaving overall utility largely intact while still enabling significant persona manipulation. While current guardrails offer partial protection, they remain brittle under sustained attack. Our findings expose new vulnerabilities in personas and highlight the need for context-resilient persona in LLMs. Our codebase and dataset is available at: https://github.com/Jivnesh/PHISH
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。