arXiv:2606.24388cs.AIcs.LG2026-06

构建大规模多模态对抗攻击数据集,助力视觉语言模型安全评估

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

  • 收集47,524个预生成对抗样本,覆盖10大类55子类有害意图
  • 整合7,826个真实意图,新增类别以提升覆盖全面性
  • 开源资源降低研究门槛,支持模型鲁棒性与防御机制测试

我们提出一个大规模、开源的视觉语言模型(VLMs)对抗攻击数据集。该数据集设计多样、具有代表性且实用,扩展了现有基准,涵盖10个高层次类别和55个子类别,覆盖多种有害意图。由于生成大量对抗样本计算成本高、流程复杂,本工作旨在使对抗数据更易获取。数据集包含47,524个由近期先进攻击策略生成的对抗样本。通过整合多个权威来源的基准并加以扩展,共收录7,826个具体意图,并新增一类以扩大覆盖范围。该数据集为研究模型鲁棒性与对齐性提供了真实评估资源。目标是让研究人员和从业者能系统评估VLM的鲁棒性与安全性,微调攻击生成模型,或在多样化对抗条件下开发和压力测试防御机制。通过发布此资源,我们希望降低对抗研究门槛,推动更具可复现性、全面性和可比性的VLM安全评估。

原文摘要 · Abstract (English)

We introduce a large-scale, open-source dataset of pre-generated adversarial attacks for vision-language models (VLMs). The dataset is designed to be diverse, representative, and practical, extending existing benchmarks by covering 10 high-level categories and 55 subcategories of harmful intents. Our primary goal is to make adversarial data accessible to the research community, given the computational cost and complexity of generating large numbers of attacks. The dataset comprises 47 524 adversarial samples, generated using state-of-the-art attack strategies from recent literature. Our work complements existing efforts by consolidating and extending prior benchmarks from multiple established sources, resulting in 7 826 intents, and introduce an additional category to broaden coverage. This provides realistic evaluation resources for studying model robustness and alignment. Our dataset intends to enable researchers and practitioners to systematically evaluate the robustness and safety of VLMs, fine-tune attack-generation models, and develop or stress-test defensive guardrails under diverse adversarial conditions. By releasing this resource, we aim to lower the barrier to adversarial research and foster more reproducible, comprehensive, and comparable evaluations of VLM safety.

对抗攻击视觉语言模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。