用醉酒语言测试大模型安全,发现易被攻破且泄露隐私
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
- 通过角色提示、因果微调和强化后训练诱导模型产生醉酒文本
- 在5个大模型上测试,攻击成功率高于基线模型及已有方法
- 方法简单高效,适合用于检测和提升模型安全性
人类饮酒后易出现不当行为和隐私泄露。本文将醉酒语言(饮酒状态下撰写的文本)作为大语言模型(LLMs)安全失效的诱因进行研究。通过三种方法诱导模型生成醉酒语言:基于角色的提示、因果微调和基于强化的学习后训练。在5个LLM上评估发现,相较于基础模型及以往方法,在JailbreakBench(含防御机制)和ConfAIde两个英文基准上,模型对越狱攻击的敏感性显著升高,且存在隐私泄露风险。通过人工评估与LLM辅助评估结合,并分析错误类别,结果表明人类醉酒行为与模型中的人格化表现存在对应关系。该方法简便高效,可作为大模型安全调优的潜在反制手段,揭示了严重的安全风险。
原文摘要 · Abstract (English)
Humans are susceptible to undesirable behaviours and privacy leaks under the influence of alcohol. This paper investigates drunk language, i.e., text written under the influence of alcohol, as a driver for safety failures in large language models (LLMs). We investigate three mechanisms for inducing drunk language in LLMs: persona-based prompting, causal fine-tuning, and reinforcement-based post-training. When evaluated on 5 LLMs, we observe a higher susceptibility to jailbreaking on JailbreakBench (even in the presence of defences) and privacy leaks on ConfAIde, where both benchmarks are in English, as compared to the base LLMs as well as previously reported approaches. Via a robust combination of manual evaluation and LLM-based evaluators and analysis of error categories, our findings highlight a correspondence between human-intoxicated behaviour, and anthropomorphism in LLMs induced with drunk language. The simplicity and efficiency of our drunk language inducement approaches position them as potential counters for LLM safety tuning, highlighting significant risks to LLM safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。