自动合成难例提升多模态模型安全检测能力
Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation

- 用多智能体迭代生成对抗性样本,无需人工标注
- 在公开数据集上将漏检率从41.2%降至24.5%
- 适合做内容安全防护的开发者和研究者
多模态大语言模型在内容安全与审核任务中应用日益广泛,但仍易受对抗攻击和分布外边缘案例影响。传统主动学习与人工标注难以应对新型多模态威胁的复杂性与规模。本文提出一种自动化、代理驱动的红队测试框架,通过迭代策略生成新假设并变异历史尝试,系统性合成难例。该框架采用多智能体架构,包括高推理能力的Architect代理、先进图像生成器及多层级验证委员会(由LLM评审员组成),可自主发现边界违规行为与模糊政策边缘案例,全程无须人工干预。通过将这些精心合成的对抗样本作为测试时检索的上下文示例,显著提升目标模型鲁棒性,在公开图像安全基准上将假阴性率(FNR)从41.2%降低至24.5%,且不依赖任何人工标注。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnerable to adversarial attacks and out-of-distribution edge cases. Traditional active learning and manual annotation fail to scale against the complexity and volume of novel multimodal threats. In this paper, we propose an automated, agentic red-teaming framework that systematically synthesizes difficult examples using an iterative strategy that proposes novel hypotheses as well as mutating on past attempts. Leveraging a multi-agent architecture that consists of a high-reasoning Architect agent, an advanced image generator, and a multi-level verification committee of LLM raters, our system autonomously uncovers boundary-pushing violations and ambiguous policy edge cases without any human intervention. By employing these carefully synthesized adversarial examples as in-context demonstrations via test-time Retrieval, we substantially improve the target model's robustness, reducing the False Negative Rate (FNR) from 41.2% to 24.5% in a public image safety benchmark without relying on any human labeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。