arXiv:2603.10010cs.CLcs.AI2026-03被引 1

FERRET通过三层扩展机制,自动生成更有效的多模态对抗对话。

FERRET: Framework for Expansion Reliant Red Teaming

  • 分三步扩展:自改进对话起点、拓展为多模态对话、动态发现攻击策略
  • 相比现有方法,在生成对抗性对话上效果更优,突破性强
  • 适合安全测试人员和模型鲁棒性研究者使用

我们提出一种多维度自动化红队测试框架 FERRET(Expansion Reliant Red Teaming),旨在生成能够破坏目标模型的多模态对抗性对话,并通过三种扩展机制提升其有效性与效率。第一,水平扩展:红队模型自我优化,生成更具针对性的对话起点;第二,垂直扩展:将已发现的对话起点拓展为高效的多模态对话;第三,元扩展:在对话过程中动态发现更有效的多模态攻击策略。实验表明,FERRET 在生成对抗性对话方面显著优于现有最先进方法。

原文摘要 · Abstract (English)

We introduce a multi-faceted automated red teaming framework in which the goal is to generate multi-modal adversarial conversations that would break a target model and introduce various expansions that would result in more effective and efficient adversarial conversations. The introduced expansions include: 1. Horizontal expansion in which the goal is for the red team model to self-improve and generate more effective conversation starters that would shape a conversation. 2. Vertical expansion in which the goal is to take these conversation starters that are discovered in the horizontal expansion phase and expand them into effective multi-modal conversations and 3. Meta expansion in which the goal is for the red team model to discover more effective multi-modal attack strategies during the course of a conversation. We call our framework FERRET (Framework for Expansion Reliant Red Teaming) and compare it with various existing automated red teaming approaches. In our experiments, we demonstrate the effectiveness of FERRET in generating effective multi-modal adversarial conversations and its superior performance against existing state of the art approaches.

红队测试对抗攻击多模态自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。