arXiv:2504.14985cs.CRcs.AI2025-04被引 3

aiXamine一站式评估大模型安全,发现主流模型存在漏洞。

aiXamine: Simplified LLM Safety and Security

  • 整合40多个测试,覆盖安全与隐私等8个维度
  • 检测出GPT-4o易受对抗攻击,Grok-3有偏见输出
  • 开源模型在公平性和鲁棒性上可媲美闭源模型

评估大型语言模型(LLMs)的安全与可靠性仍是一项复杂任务,常需用户面对零散的基准测试、数据集、指标和报告格式。为解决此问题,我们提出aiXamine,一个全面的黑盒评估平台,用于评估LLM的安全与安全。aiXamine集成超过40项测试(即基准),分为八大服务,涵盖对抗鲁棒性、代码安全、公平性与偏见、幻觉、模型与数据隐私、分布外(OOD)鲁棒性、过度拒绝及安全对齐。该平台将评估结果汇总为每模型一份详细报告,包含性能分解、测试样例与丰富可视化。我们使用aiXamine评估了50多个公开及专有LLM,共执行2000余次检查。结果揭示主流模型存在显著缺陷:OpenAI GPT-4o易受对抗攻击,xAI Grok-3存在输出偏见,Google Gemini 2.0存在隐私弱点。此外,我们发现开源模型在安全对齐、公平性与偏见、以及OOD鲁棒性方面可达到甚至超越闭源模型表现。最后,我们识别出蒸馏策略、模型规模、训练方法与架构选择之间的权衡关系。

原文摘要 · Abstract (English)

Evaluating Large Language Models (LLMs) for safety and security remains a complex task, often requiring users to navigate a fragmented landscape of ad hoc benchmarks, datasets, metrics, and reporting formats. To address this challenge, we present aiXamine, a comprehensive black-box evaluation platform for LLM safety and security. aiXamine integrates over 40 tests (i.e., benchmarks) organized into eight key services targeting specific dimensions of safety and security: adversarial robustness, code security, fairness and bias, hallucination, model and data privacy, out-of-distribution (OOD) robustness, over-refusal, and safety alignment. The platform aggregates the evaluation results into a single detailed report per model, providing a detailed breakdown of model performance, test examples, and rich visualizations. We used aiXamine to assess over 50 publicly available and proprietary LLMs, conducting over 2K examinations. Our findings reveal notable vulnerabilities in leading models, including susceptibility to adversarial attacks in OpenAI's GPT-4o, biased outputs in xAI's Grok-3, and privacy weaknesses in Google's Gemini 2.0. Additionally, we observe that open-source models can match or exceed proprietary models in specific services such as safety alignment, fairness and bias, and OOD robustness. Finally, we identify trade-offs between distillation strategies, model size, training methods, and architectural choices.

大模型安全评估平台AI审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。