arXiv:2601.03273cs.CLcs.AI2026-01被引 1

构建多视角安全评测集,提升大模型对隐性有害内容的识别能力。

A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness

  • 设计涵盖106类细粒度的安全评测数据集GuardEval
  • GGuard模型在复杂边缘案例上达到0.832的宏平均F1分数
  • 适合需要高精度内容审核与抗攻击能力的研究者使用

随着大语言模型深度融入日常生活,亟需更安全的内容审核系统,以区分无意请求与有害内容,并保持恰当的审查边界。现有模型在处理隐性冒犯、微妙性别与种族偏见及越狱提示等复杂情况时表现不佳,因其依赖训练数据而可能强化社会偏见,导致输出不一致且伦理问题频发。为此,我们提出GuardEval,一个统一的多视角基准数据集,包含106个细粒度类别,覆盖人类情绪、攻击性语言、性别与种族偏见及广泛安全关切。同时,我们推出基于Gemma3-12B的量化低秩适配(QLoRA)模型GemmaGuard(GGuard),在GuardEval上进行微调,用于细粒度内容审核。评估显示,GGuard在宏平均F1上达0.832,显著优于OpenAI Moderator(0.64)和Llama Guard(0.61)。研究证明,多视角、以人为本的安全基准对减少审核不一致至关重要。GuardEval与GGuard共同表明,多样化、代表性数据能实质性提升安全性和对抗鲁棒性。

原文摘要 · Abstract (English)

As large language models (LLMs) become deeply embedded in daily life, the urgent need for safer moderation systems that distinguish between naive and harmful requests while upholding appropriate censorship boundaries has never been greater. While existing LLMs can detect dangerous or unsafe content, they often struggle with nuanced cases such as implicit offensiveness, subtle gender and racial biases, and jailbreak prompts, due to the subjective and context-dependent nature of these issues. Furthermore, their heavy reliance on training data can reinforce societal biases, resulting in inconsistent and ethically problematic outputs. To address these challenges, we introduce GuardEval, a unified multi-perspective benchmark dataset designed for both training and evaluation, containing 106 fine-grained categories spanning human emotions, offensive and hateful language, gender and racial bias, and broader safety concerns. We also present GemmaGuard (GGuard), a Quantized Low-Rank Adaptation (QLoRA), fine-tuned version of Gemma3-12B trained on GuardEval, to assess content moderation with fine-grained labels. Our evaluation shows that GGuard achieves a macro F1 score of 0.832, substantially outperforming leading moderation models, including OpenAI Moderator (0.64) and Llama Guard (0.61). We show that multi-perspective, human-centered safety benchmarks are critical for mitigating inconsistent moderation decisions. GuardEval and GGuard together demonstrate that diverse, representative data materially improve safety, and adversarial robustness on complex, borderline cases.

内容安全多视角评测大模型防御偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。