arXiv:2502.01236cs.LGcs.AI2025-02ICML被引 24

用智能代理自动挖掘让大模型产生幻觉或越狱的提示词。

Eliciting Language Model Behaviors with Investigator Agents

  • 训练调查代理模型,通过迭代优化生成多样提示以激发目标行为。
  • 在AdvBench数据集上实现100%越狱成功率,幻觉率高达85%。
  • 生成的提示可解释性强,适用于安全测试与对抗攻击研究。

语言模型在自由文本提示下表现出复杂多样的行为,难以全面描述其输出空间。本文研究行为诱发问题,旨在搜索能触发特定目标行为(如幻觉或有害响应)的提示。为应对提示空间指数级增长的挑战,我们训练调查代理模型,将随机选择的目标行为映射到能诱发该行为的多样化输出分布,类似近似贝叶斯推断。通过监督微调、基于DPO的强化学习以及一种新颖的Frank-Wolfe训练目标,迭代发现多样化的提示策略。实验表明,该方法能生成多种有效且人类可解释的提示,成功引发越狱、幻觉及开放性异常行为,在AdvBench(有害行为子集)上达到100%攻击成功率,幻觉率达85%。

原文摘要 · Abstract (English)

Language models exhibit complex, diverse behaviors when prompted with free-form text, making it difficult to characterize the space of possible outputs. We study the problem of behavior elicitation, where the goal is to search for prompts that induce specific target behaviors (e.g., hallucinations or harmful responses) from a target language model. To navigate the exponentially large space of possible prompts, we train investigator models to map randomly-chosen target behaviors to a diverse distribution of outputs that elicit them, similar to amortized Bayesian inference. We do this through supervised fine-tuning, reinforcement learning via DPO, and a novel Frank-Wolfe training objective to iteratively discover diverse prompting strategies. Our investigator models surface a variety of effective and human-interpretable prompts leading to jailbreaks, hallucinations, and open-ended aberrant behaviors, obtaining a 100% attack success rate on a subset of AdvBench (Harmful Behaviors) and an 85% hallucination rate.

提示工程安全测试幻觉检测对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。