arXiv:2507.14719cs.AI2025-07被引 1

用政策生成对抗性测试,评估20个大模型在10个安全领域的表现差异。

Policy-Grounded Safety Evaluation of 20 Large Language Models

  • 将安全政策转为对抗性提示,用AI评分器评估响应安全性。
  • 平均安全分从86.2%到52.4%,隐私与冒充领域仅24.3%。
  • 适合关注模型安全评估、政策落地的开发者和监管者。

随着大语言模型(LLMs)日益融入真实应用场景,可扩展且严谨的安全评估至关重要。本文提出Aymara AI,一个用于生成和管理定制化、基于政策的安全评估的程序化平台。Aymara AI将自然语言安全政策转化为对抗性提示,并使用经过人类判断验证的AI评分器对模型响应进行打分。通过Aymara LLM风险与责任矩阵,我们评估了20个商用大模型在10个真实世界安全领域的表现。结果显示性能差异显著,平均安全得分介于86.2%至52.4%之间。尽管模型在误导信息等成熟安全领域表现良好(均值95.7%),但在更复杂或定义不清的领域持续失败,尤其是隐私与冒充领域(均值24.3%)。方差分析证实,模型和领域间的安全得分存在显著差异(p < .05)。这些发现凸显了大模型安全性的不一致性和情境依赖性,强调了像Aymara AI这类可扩展、可定制工具在推动负责任人工智能发展与监督中的必要性。

原文摘要 · Abstract (English)

As large language models (LLMs) become increasingly integrated into real-world applications, scalable and rigorous safety evaluation is essential. This paper introduces Aymara AI, a programmatic platform for generating and administering customized, policy-grounded safety evaluations. Aymara AI transforms natural-language safety policies into adversarial prompts and scores model responses using an AI-based rater validated against human judgments. We demonstrate its capabilities through the Aymara LLM Risk and Responsibility Matrix, which evaluates 20 commercially available LLMs across 10 real-world safety domains. Results reveal wide performance disparities, with mean safety scores ranging from 86.2% to 52.4%. While models performed well in well-established safety domains such as Misinformation (mean = 95.7%), they consistently failed in more complex or underspecified domains, notably Privacy & Impersonation (mean = 24.3%). Analyses of Variance confirmed that safety scores differed significantly across both models and domains (p < .05). These findings underscore the inconsistent and context-dependent nature of LLM safety and highlight the need for scalable, customizable tools like Aymara AI to support responsible AI development and oversight.

大模型安全评估框架政策落地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。