arXiv:2505.11365cs.CYcs.AI2025-05被引 1

Phare框架可系统检测大模型在幻觉、偏见和有害内容上的安全漏洞。

Phare: A Safety Probe for Large Language Models

  • 构建多语言诊断框架,从三方面评估模型安全性
  • 发现17个顶尖模型普遍存在顺从、提示敏感等系统性缺陷
  • 聚焦具体失效模式,帮助开发者改进模型可靠性

确保大语言模型(LLMs)的安全性对负责任部署至关重要,但现有评估往往侧重性能而忽视故障模式。我们提出Phare,一个跨语言的诊断框架,用于探测和评估LLM在幻觉与可靠性、社会偏见、有害内容生成三个关键维度的表现。对17个最先进的LLMs进行评估后发现,所有安全维度均存在系统性脆弱性,包括迎合倾向、提示敏感性和刻板印象再现。Phare不简单排名模型,而是揭示具体失效模式,为研究人员和从业者提供可操作的洞察,以构建更稳健、对齐且可信的语言系统。

原文摘要 · Abstract (English)

Ensuring the safety of large language models (LLMs) is critical for responsible deployment, yet existing evaluations often prioritize performance over identifying failure modes. We introduce Phare, a multilingual diagnostic framework to probe and evaluate LLM behavior across three critical dimensions: hallucination and reliability, social biases, and harmful content generation. Our evaluation of 17 state-of-the-art LLMs reveals patterns of systematic vulnerabilities across all safety dimensions, including sycophancy, prompt sensitivity, and stereotype reproduction. By highlighting these specific failure modes rather than simply ranking models, Phare provides researchers and practitioners with actionable insights to build more robust, aligned, and trustworthy language systems.

模型安全大模型评估偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。