arXiv:2608.10171cs.AIcs.CR2026-08

用生成流网络自动攻破大模型,还能生成土耳其语攻击

Generating Attacks for LLMs with GFlowNets

论文配图:Generating Attacks for LLMs with GFlowNets
图 1 · 摘自论文原文
  • 用GFlowNets训练攻击模型,自动对抗目标大模型
  • 英文攻击效果优于现有基准,首次实现土耳其语攻击生成
  • 无需人工干预,可量化评估模型鲁棒性,适合安全测试

大型语言模型(LLMs)的快速发展使其在多个领域广泛应用,但也带来了显著的安全漏洞,亟需识别并缓解恶意利用引发的缺陷。红队测试通过多样化的对抗输入评估模型鲁棒性,是发现安全风险与实施防御措施的关键。当前红队测试或依赖专家手动执行,或使用预定义攻击数据集自动化进行。然而,人工测试耗时长,现有自动化方法因依赖固定数据集而创造力有限。本研究提出一种无需人工参与、可自适应的自动化方法,利用生成流网络(GFlowNets)实现一个大模型对另一个大模型的测试。在此框架中,攻击模型针对特定目标模型进行训练,以实现自动化红队测试并提供量化鲁棒性评分。本研究旨在生成比现有基准更有效的英文对抗攻击,并作为文献中的新贡献,首次引入可在土耳其语生成攻击输入的模型。

原文摘要 · Abstract (English)

The rapid advancement of Large Language Models (LLMs) has facilitated their ubiquitous integration into various domains, leading to widespread adoption. However, this escalating trend has introduced significant security vulnerabilities, necessitating the identification and mitigation of flaws arising from malicious exploitation. Red teaming assessments, conducted to evaluate model robustness through diverse adversarial inputs, are essential for exposing security risks and implementing countermeasures. Currently, red teaming is performed either manually by experts or automatically using predefined attack datasets. Nevertheless, manual testing remains time-consuming, while existing automated methods suffer from limited creativity due to their inherent dependency on fixed datasets. In this study, we propose an automated, human-independent, and adaptive approach leveraging GFlowNets to identify LLM vulnerabilities by utilizing one large language model to test another. Within this framework, an attacker model is trained against a specified victim model to perform automated red teaming and provide a quantitative robustness score. This research aims to generate more effective adversarial attacks in English compared to existing benchmarks and, as a novel contribution to the literature, introduces a model capable of generating attack inputs in the Turkish language.

大模型安全对抗攻击GFlowNets红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。