用心理学方法设计多轮对话攻击,突破大模型安全防线。
PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies

- 基于社会心理学构建多轮说服策略,精准操控模型响应
- 在4个对齐模型上实现87.3%平均成功率,显著优于现有方法
- 发现模型特异性心理指纹,揭示不同模型的脆弱性差异
大型语言模型(LLMs)正越来越多地应用于教育、医疗、政策咨询等需持续交互的场景,用户将其视为长期对话伙伴而非单次查询工具。这种转变使多轮对话攻击成为日益严重的安全威胁,但现有研究多集中于单轮提示优化或迭代攻击改进,缺乏对心理机制驱动的多轮漏洞探索。本文提出PsychJail,一个基于心理学理论的红队框架,通过理论指导的多轮说服来测试对齐模型的安全性。该框架将已知的社会心理说服技术映射为策略条件化的攻击策略,将每次攻击动作分解为意义变更分析、策略选择和可见消息生成,实现说服知识模型(PKM)的可操作化。策略通过轨迹级强化学习进行优化,采用受PKM约束的奖励函数,仅当每一轮均包含完整的意义变更分析时才给予早期成功奖励。在四个对齐的受害者模型上,PsychJail实现了87.3%的平均攻击成功率,超越所有强基线。我们还测量了每个模型被攻破的动作,揭示四种模型级指纹,识别出影响各模型的说服杠杆及其作用范围。这些指纹有助于解释跨模型迁移的不对称性。我们将其解释为四种候选心理画像:理性主义型、可信度驱动型、叙事单一型与普遍易说服型,但此解释仍为假设,有待未来验证。研究确立了心理型越狱攻击作为日益交互式大模型红队测试的新前沿。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage them as sustained social interlocutors rather than one-shot query engines. This shift makes jailbreaks a growing safety threat, yet most research emphasizes single-turn prompt optimization or iterative attack refinement, leaving psychologically grounded multi-turn vulnerabilities underexplored. We present PsychJail, a psychology-guided framework for red teaming aligned LLMs through theory-grounded, multi-turn persuasion. PsychJail maps established social-psychological persuasion techniques into a tactic-conditioned attack policy. It factorizes each attacker action into a Change-of-Meaning analysis, tactic selection, and victim-visible message, operationalizing the Persuasion Knowledge Model (PKM). The policy is refined with trajectory-level reinforcement learning using a PKM-gated reward that credits early jailbreak success only when every turn contains a well-formed Change-of-Meaning analysis. Across four aligned victim models, PsychJail achieves the highest average attack success rate (87.3%) and outperforms strong single-turn and multi-turn baselines on every model. We also measure susceptibility at the action that breaks each victim, revealing four distinct model-level fingerprints that identify which persuasion levers affect each model and how broadly. These fingerprints help explain cross-model transfer asymmetry. We interpret them as four candidate psychological profiles-rationalist, credibility-driven, narrative-monoculture, and broadly persuadable-while treating this interpretation as a conjecture requiring future validation. Our findings establish psychological jailbreaks as a distinct red-teaming frontier for increasingly interactive LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。