用道德攻击破解大模型,暴露其价值观脆弱性。
Jailbreaking Large Language Models with Morality Attacks

- 构建10.3K条道德模糊与冲突数据集,设计四类对抗攻击
- 大模型和安全护栏模型均易被微妙道德攻击误导
- 揭示多价值观对齐中的安全漏洞,适合安全研究者参考
多元价值观对齐是使AI与多元人类共存并服务的重要目标。尽管已有大量研究致力于提升大语言模型(LLMs)的多元价值观学习能力,但其在多元价值下生成道德内容的鲁棒性仍待深入探索。受越狱提示惊人说服力启发,我们提出利用越狱攻击来研究大模型内在的多元价值观。具体地,我们构建了一个包含10.3K条样本的道德数据集,分为价值模糊与价值冲突两类,并据此形式化出四种对抗攻击,以操纵大模型对道德问题的判断。我们在具备灵活用户输入能力的大语言模型及典型生成系统中的护栏模型上进行评估。实验结果表明,大模型和护栏模型对这些细微且复杂的道德感知攻击存在显著脆弱性。
原文摘要 · Abstract (English)
Pluralism alignment with AI has the sophisticated and necessary goal of creating AI that can coexist with and serve morally multifaceted humanity. Research towards pluralism alignment has many efforts in enhancing the learning of large language models (LLMs) to accomplish pluralism. Although this is essential, the robustness of LLMs to produce moral content over pluralistic values is still under exploration.Inspired by the astonishing persuasion abilities via jailbreak prompts, we propose to leverage jailbreak attacks to study LLMs' internal pluralistic values. In detail, we develop a morality dataset with 10.3K instances in two categories: Value Ambiguity and Value Conflict. We further formalize four adversarial attacks with the constructed dataset, to manipulate LLMs' judgment over the morality questions. We evaluate both the large language models and guardrail models which are typically used in generative systems with flexible user input. Our experiment results show that there is a critical vulnerability of LLMs and guardrail models to these subtle and sophisticated moral-aware attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。