用自动化流程四天建成低冗余高区分度的LLM安全测试集
SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
- 七类智能代理协同生成安全测试题,全程无需人工干预
- 构建出23,446条低冗余查询数据,可有效区分模型安全性能
- 适合需要高效评估大模型安全性的研究者与开发者使用
大语言模型(LLMs)的快速普及加剧了对可靠安全评估的需求,以发现模型漏洞。然而,现有安全评测基准多依赖耗时费力的人工构建,存在资源消耗大、重复度高、难度有限等问题。为此,我们提出SafetyFlow,首个用于自动化构建LLM安全评测基准的代理流系统。SafetyFlow通过协调七个专业代理,在无任何人工干预的情况下仅用四天即可完成全面的安全基准构建,显著降低时间和资源成本。系统配备多样化工具,确保流程与成本可控,并将人类专家经验融入自动化流程。最终构建的数据集SafetyFlowBench包含23,446条查询,冗余低且具有强区分能力。我们评估了49个先进LLM在该数据集上的安全性,并通过大量实验验证了其有效性与效率。
原文摘要 · Abstract (English)
The rapid proliferation of large language models (LLMs) has intensified the requirement for reliable safety evaluation to uncover model vulnerabilities. To this end, numerous LLM safety evaluation benchmarks are proposed. However, existing benchmarks generally rely on labor-intensive manual curation, which causes excessive time and resource consumption. They also exhibit significant redundancy and limited difficulty. To alleviate these problems, we introduce SafetyFlow, the first agent-flow system designed to automate the construction of LLM safety benchmarks. SafetyFlow can automatically build a comprehensive safety benchmark in only four days without any human intervention by orchestrating seven specialized agents, significantly reducing time and resource cost. Equipped with versatile tools, the agents of SafetyFlow ensure process and cost controllability while integrating human expertise into the automatic pipeline. The final constructed dataset, SafetyFlowBench, contains 23,446 queries with low redundancy and strong discriminative power. Our contribution includes the first fully automated benchmarking pipeline and a comprehensive safety benchmark. We evaluate the safety of 49 advanced LLMs on our dataset and conduct extensive experiments to validate our efficacy and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。