打造统一工具箱,让大模型安全研究更可复现、可比较。
AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
- 整合12种攻击算法与7个基准数据集,支持多维度评测
- 支持确定性结果与资源追踪,确保实验可复现
- 适合作为研究者开展大模型安全实验的标准化平台
大型语言模型(LLM)安全与鲁棒性研究快速发展,但相关实现、数据集和评估方法碎片化且常含错误,导致研究难以复现和比较。为此,我们提出AdversariaLLM,一个用于大模型越狱鲁棒性研究的统一工具箱。该框架强调可复现性、正确性和可扩展性,实现12种对抗攻击算法,集成7个涵盖有害性、过度拒绝和实用性评估的基准数据集,并通过Hugging Face提供多种开源权重模型。其支持计算资源追踪、确定性结果输出及分布评估技术,提升实验可比性。配套组件JudgeZoo可独立用于评判,共同构建透明、可比较、可复现的大模型安全研究基础。
原文摘要 · Abstract (English)
The rapid expansion of research on Large Language Model (LLM) safety and robustness has produced a fragmented and oftentimes buggy ecosystem of implementations, datasets, and evaluation methods. This fragmentation makes reproducibility and comparability across studies challenging, hindering meaningful progress. To address these issues, we introduce AdversariaLLM, a toolbox for conducting LLM jailbreak robustness research. Its design centers on reproducibility, correctness, and extensibility. The framework implements twelve adversarial attack algorithms, integrates seven benchmark datasets spanning harmfulness, over-refusal, and utility evaluation, and provides access to a wide range of open-weight LLMs via Hugging Face. The implementation includes advanced features for comparability and reproducibility such as compute-resource tracking, deterministic results, and distributional evaluation techniques. \name also integrates judging through the companion package JudgeZoo, which can also be used independently. Together, these components aim to establish a robust foundation for transparent, comparable, and reproducible research in LLM safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。