构建了前沿大模型有害输出的基准数据集,揭示其风险特征。
HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

- 以内容为中心构建跨模型家族的有害输出数据集
- 23个前沿模型产生超8万条验证过的有害内容,15类危害
- 模型越强,有害内容越多且类型越多样,表面安全下藏深层风险
前沿大语言模型的安全评估长期将有害生成视为攻击结果而非分析对象,导致对其错误行为产出的有害内容了解甚少,主要因大规模高质量的模型误用数据难获取。为此,我们提出HarmProfile,一个以内容为核心的基准数据集,涵盖多种危害类别和模型家族的模型误用行为,并将有害输出分布定义为模型级风险画像。核心思想是:如同从语料库可刻画语言行为,也可通过有害内容、严重程度与变异度刻画模型风险。HarmProfile包含来自23个前沿大模型、13个模型家族的超8万条经验证的有害样本,划分为15类危害与57个子类。基于该数据集,我们发现前沿模型在规模上持续生成有害内容,但各模型风险特征各异;有害性与多样性随模型能力提升而增长,表明模型可能表面安全,却潜藏日益危险的知识储备。代码已开源。
原文摘要 · Abstract (English)
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmful outputs produced during model misbehavior, partly because large-scale, high-quality collections of frontier-LLM misbehavior are difficult to obtain. To address this gap, we introduce HarmProfile, a content-centric benchmark dataset that collects model misbehavior across diverse harm categories and model families, and defines the resulting harmful-output distribution as a model-level risk profile. The premise is that, just as linguistic behavior can be characterized from an utterance corpus, model risk can be characterized from the content, severity, and variation of its safety failures. HarmProfile contains over 80,000 validated artifacts from 23 frontier LLMs across 13 model families, organized into 15 harm categories and 57 subcategories. Using this corpus, we find that frontier LLMs reliably produce harmful content at scale, yet exhibit distinct risk profiles; both harmfulness and diversity grow with model capability, suggesting that frontier LLMs may appear safe yet harbor increasingly dangerous knowledge beneath the alignment surface. Our source code is available at https://github.com/fresh-ma/HarmProfile .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。