arXiv:2409.04822cs.CLcs.AI2024-09被引 3

用现成大模型做对话红队测试,能自适应攻击策略。

Exploring Straightforward Conversational Red-Teaming

  • 用现成LLM模拟攻击者,通过多轮对话诱使目标模型生成不当输出。
  • 模型能根据之前尝试调整策略,但对齐度越高效果越差。
  • 适合安全测试人员评估模型对抗风险。

大型语言模型(LLMs)在商业对话系统中应用日益广泛,但存在安全与伦理风险。多轮对话中,上下文会影响模型行为,可能被利用以生成不良响应。本文研究了使用现成的LLM进行直接红队测试的有效性,攻击者LLM旨在诱使目标LLM产生不良输出,比较单轮与对话式红队策略。实验揭示了多种使用策略对红队表现有显著影响。结果表明,现成模型可作为有效的红队工具,并能根据过往尝试调整攻击策略,尽管其有效性随模型对齐程度提升而下降。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used in business dialogue systems but they pose security and ethical risks. Multi-turn conversations, where context influences the model's behavior, can be exploited to produce undesired responses. In this paper, we examine the effectiveness of utilizing off-the-shelf LLMs in straightforward red-teaming approaches, where an attacker LLM aims to elicit undesired output from a target LLM, comparing both single-turn and conversational red-teaming tactics. Our experiments offer insights into various usage strategies that significantly affect their performance as red teamers. They suggest that off-the-shelf models can act as effective red teamers and even adjust their attack strategy based on past attempts, although their effectiveness decreases with greater alignment.

红队测试LLM安全对话攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。