arXiv:2508.07063cs.AI2025-08被引 5

构建统一数据集评估大模型内容审核能力,发现其在隐性偏见检测上仍有明显不足。

Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach

  • 构建涵盖49类情感与偏见的统一评测数据集
  • 自研SafePhi模型在宏观F1上达0.89,优于OpenAI和Llama Guard
  • 强调需引入人类参与以提升模型公平性与可解释性

随着AI系统日益融入日常生活,安全可靠的审核机制需求愈发迫切。尽管大语言模型(LLMs)在复杂任务中表现出色,但在涉及微妙道德判断的领域仍易出错,难以识别隐性仇恨、攻击性语言及性别偏见。其输出常受训练数据影响,加剧社会偏见。为此,本文构建了一个包含49个类别、覆盖广泛情绪、攻击性文本与性别种族偏见的统一基准数据集。基于SOTA模型,提出SafePhi——一个经过QLoRA微调的Phi-4版本,在多类别分类中实现0.89的宏观F1得分,显著优于OpenAI Moderator(0.77)和Llama Guard(0.74)。研究同时揭示了模型在若干关键领域的持续短板,呼吁采用更异质、更具代表性的数据,并引入人机协同机制,以增强模型鲁棒性与可解释性。

原文摘要 · Abstract (English)

As AI systems become more integrated into daily life, the need for safer and more reliable moderation has never been greater. Large Language Models (LLMs) have demonstrated remarkable capabilities, surpassing earlier models in complexity and performance. Their evaluation across diverse tasks has consistently showcased their potential, enabling the development of adaptive and personalized agents. However, despite these advancements, LLMs remain prone to errors, particularly in areas requiring nuanced moral reasoning. They struggle with detecting implicit hate, offensive language, and gender biases due to the subjective and context-dependent nature of these issues. Moreover, their reliance on training data can inadvertently reinforce societal biases, leading to inconsistencies and ethical concerns in their outputs. To explore the limitations of LLMs in this role, we developed an experimental framework based on state-of-the-art (SOTA) models to assess human emotions and offensive behaviors. The framework introduces a unified benchmark dataset encompassing 49 distinct categories spanning the wide spectrum of human emotions, offensive and hateful text, and gender and racial biases. Furthermore, we introduced SafePhi, a QLoRA fine-tuned version of Phi-4, adapting diverse ethical contexts and outperforming benchmark moderators by achieving a Macro F1 score of 0.89, where OpenAI Moderator and Llama Guard score 0.77 and 0.74, respectively. This research also highlights the critical domains where LLM moderators consistently underperformed, pressing the need to incorporate more heterogeneous and representative data with human-in-the-loop, for better model robustness and explainability.

内容审核大模型评估伦理对齐人类在环

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。