arXiv:2506.00973cs.CL2025-06ACL被引 5

为大模型极端内容生成安全风险设计分级评估体系

XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content

  • 构建包含3840个真实场景提示的分级评测基准
  • 将安全失效分为5个危险等级,量化风险严重程度
  • 提出可解释的攻击严重度曲线,适合安全研究者使用

大型语言模型(LLMs)可能生成从意识形态言论到暴力指导的各种内容。现有安全评估多依赖简单的二元标签(安全/不安全),忽略了输出风险的复杂谱系。为此,我们提出XGUARD,一个用于评估LLMs生成极端内容严重性的基准与评估框架。XGUARD包含3,840个来自社交媒体、新闻等真实数据的红队提示,覆盖广泛意识形态情境。框架将模型输出划分为五个危险等级(0至4),支持对失败频率与严重性的精细分析。我们引入可解释的攻击严重度曲线(ASC),可视化模型在不同威胁强度下的脆弱性,并比较多种防御机制。通过XGUARD评估六款主流LLM及两种轻量级防御策略,揭示了当前安全短板与鲁棒性与表达自由间的权衡。本工作凸显了分级安全度量对构建可信LLMs的重要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) can generate content spanning ideological rhetoric to explicit instructions for violence. However, existing safety evaluations often rely on simplistic binary labels (safe and unsafe), overlooking the nuanced spectrum of risk these outputs pose. To address this, we present XGUARD, a benchmark and evaluation framework designed to assess the severity of extremist content generated by LLMs. XGUARD includes 3,840 red teaming prompts sourced from real world data such as social media and news, covering a broad range of ideologically charged scenarios. Our framework categorizes model responses into five danger levels (0 to 4), enabling a more nuanced analysis of both the frequency and severity of failures. We introduce the interpretable Attack Severity Curve (ASC) to visualize vulnerabilities and compare defense mechanisms across threat intensities. Using XGUARD, we evaluate six popular LLMs and two lightweight defense strategies, revealing key insights into current safety gaps and trade-offs between robustness and expressive freedom. Our work underscores the value of graded safety metrics for building trustworthy LLMs.

大模型安全极端内容评测基准风险分级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。