arXiv:2603.22882cs.LGcs.CV2026-03被引 3

用动态演化策略树自动挖掘视觉语言模型漏洞,效果远超传统方法。

TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy Exploration

  • 基于大模型构建策略树,动态决定探索或优化攻击路径。
  • 在12个主流模型上达成最高87.6%攻击成功率,覆盖多数公开漏洞。
  • 生成攻击更隐蔽且多样,适合安全研究人员验证模型鲁棒性。

视觉语言模型(VLM)的快速发展凸显了其安全漏洞。现有红队测试方法受限于线性探索范式,仅能在预设策略集中优化,难以发现新攻击。为此,我们提出TreeTeaming,一种将策略探索从静态测试转变为动态进化发现的新框架。核心是一个由大语言模型驱动的策略协调器,可自主决策是否演化有潜力的攻击路径或探索多样化策略分支,从而动态构建并扩展策略树。随后,多模态执行器负责实施这些复杂策略。在12个主流VLM上的实验表明,TreeTeaming在11个模型上达到当前最优攻击成功率,最高达GPT-4o的87.60%。框架还展现出超过所有公开越狱策略集合的战略多样性。此外,生成攻击平均毒性降低23.09%,体现其隐蔽性。本工作提出了自动化漏洞发现的新范式,强调必须超越静态启发式方法,以主动探索保障前沿AI模型安全。

原文摘要 · Abstract (English)

The rapid advancement of Vision-Language Models (VLMs) has brought their safety vulnerabilities into sharp focus. However, existing red teaming methods are fundamentally constrained by an inherent linear exploration paradigm, confining them to optimizing within a predefined strategy set and preventing the discovery of novel, diverse exploits. To transcend this limitation, we introduce TreeTeaming, an automated red teaming framework that reframes strategy exploration from static testing to a dynamic, evolutionary discovery process. At its core lies a strategic Orchestrator, powered by a Large Language Model (LLM), which autonomously decides whether to evolve promising attack paths or explore diverse strategic branches, thereby dynamically constructing and expanding a strategy tree. A multimodal actuator is then tasked with executing these complex strategies. In the experiments across 12 prominent VLMs, TreeTeaming achieves state-of-the-art attack success rates on 11 models, outperforming existing methods and reaching up to 87.60\% on GPT-4o. The framework also demonstrates superior strategic diversity over the union of previously public jailbreak strategies. Furthermore, the generated attacks exhibit an average toxicity reduction of 23.09\%, showcasing their stealth and subtlety. Our work introduces a new paradigm for automated vulnerability discovery, underscoring the necessity of proactive exploration beyond static heuristics to secure frontier AI models.

红队测试视觉语言模型漏洞挖掘自动化攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。