小模型实现高安全检测,多语言表现领先。
HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety
- 用46条政策构建安全宪法,生成带反事实对比的数据
- 0.8B模型在多语言评测中F1达90.9,误报率仅4.3%
- 开源模型,持续对抗测试强化防御能力
我们提出HaloGuard 1.0,一个开放权重的宪法式输入安全分类器。其在英语和多语言提示安全基准上达到顶尖性能,模型规模仅为当前领先开源防护模型的十分之一。安全宪法作为语料组织结构,包含46项政策与2,940个子类别,驱动合成数据生成:通过固定话题与词汇、仅反转意图的成对反事实样本,实现两阶段无害设计,分别处理边界与基础误报;在46种语言间均衡分布,将语言视为边界两侧的表层形式而非对抗信号。在七个提示安全基准上,HaloGuard 1.0-0.8B的平均F1为90.9,优于最大达27B参数的基线(超过30倍大),同时保持误报率(FPR)4.3%、误报率(FNR)9.5%。HaloGuard 1.0-4B版本平均F1达92.1,误报率3.5%,将额外容量用于提升精确率而非召回率。对剩余失败案例的结构化审查表明,多数看似漏检实为基准标注错误。持续运行的对抗红队协议不断强化模型应对内容级与代理型攻击的能力。模型已作为开放权重发布。
原文摘要 · Abstract (English)
We present HaloGuard 1.0, an open-weights implementation of the constitutional-classifier paradigm for input safety. It achieves state-of-the-art performance on English and multilingual prompt-safety benchmarks at roughly one-tenth the model size of current leading open guard models. The safety constitution is the organising structure of the corpus: a natural-language constitution of 46 policies and 2,940 subcategories drives synthetic data generation, with exhaustive one-to-one paired counterfactuals that hold topic and vocabulary fixed while flipping intent, a two-tier harmless design that separately targets boundary and baseline false positives (FPs), and balanced multilingual materialisation across 46 languages that treats language as a surface form appearing on both sides of the boundary rather than as an adversarial signal. Across seven prompt-safety benchmarks, HaloGuard 1.0-0.8B attains the best average F1 (90.9) of any open guard we evaluate, outperforming baselines up to 27B parameters (over 30 times larger) while holding false-positive rate (FPR) to 4.3 and false-negative rate (FNR) to 9.5. The HaloGuard 1.0-4B variant reaches average F1 of 92.1 and FPR of 3.5, spending its extra capacity on precision rather than recall. A structured adjudication of the remaining failures indicates that most apparent missed-harm cases are benchmark mislabels rather than genuine model misses. An always-on adversarial red-teaming protocol continuously hardens the guard against both content-level and agentic attacks. We release the models as open weights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。