arXiv:2601.19507cs.CL2026-01被引 1

首个自动化LVLM安全评测系统,高效生成高质安全测试集

Automated Safety Benchmarking: A Multi-agent Pipeline for LVLMs

  • 四代理协同构建:预处理、生成、增强、筛选,全流程自动化
  • 一周内完成高质量基准构建,最安全与最不安全模型安全率差达70%
  • 适合模型安全评估者、安全研究者,尤其关注快速迭代的视觉语言模型

大型视觉语言模型(LVLM)在跨模态任务中表现卓越,但面临严峻的安全挑战,影响其在真实场景中的可靠性。现有安全评测基准受限于人工成本高、复杂度静态、区分能力弱等问题,难以跟上模型快速演进和新兴风险。为此,我们提出VLSafetyBencher,首个面向LVLM安全评测的自动化系统。该系统引入四个协作智能体:数据预处理、生成、增强与选择,实现高质量样本的自动构造与筛选。实验表明,VLSafetyBencher可在一周内以极低成本构建优质安全基准,有效区分不同模型的安全性,最安全与最不安全模型间安全率差距达70%。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) exhibit remarkable capabilities in cross-modal tasks but face significant safety challenges, which undermine their reliability in real-world applications. Efforts have been made to build LVLM safety evaluation benchmarks to uncover their vulnerability. However, existing benchmarks are hindered by their labor-intensive construction process, static complexity, and limited discriminative power. Thus, they may fail to keep pace with rapidly evolving models and emerging risks. To address these limitations, we propose VLSafetyBencher, the first automated system for LVLM safety benchmarking. VLSafetyBencher introduces four collaborative agents: Data Preprocessing, Generation, Augmentation, and Selection agents to construct and select high-quality samples. Experiments validates that VLSafetyBencher can construct high-quality safety benchmarks within one week at a minimal cost. The generated benchmark effectively distinguish safety, with a safety rate disparity of 70% between the most and least safe models.

LVLM安全自动化评测多智能体基准构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。