arXiv:2505.15365cs.HCcs.CL2025-05被引 4

AI评价系统比人类更认可模型拒绝敏感请求,可能隐藏价值偏差。

AI vs. Human Judgment of Content Moderation: LLM-as-a-Judge and Ethics-Based Response Refusals

  • 用AI当裁判评估模型拒答表现,区分伦理与技术性拒绝。
  • AI裁判对伦理拒答评分显著高于人类用户,技术拒答无差异。
  • 发现评价系统存在内容审核偏见,影响AI安全训练公平性。

随着大语言模型(LLMs)在高风险场景中的广泛应用,其对涉及仇恨言论或非法活动等敏感请求的拒绝能力已成为内容审核与负责任AI实践的核心。尽管拒绝响应可体现伦理对齐与安全意识,但近期研究显示用户可能对此类响应持负面看法。与此同时,自动化评估在模型评测与训练中日益重要,其中LLM-as-a-Judge框架(即一个模型评估另一个模型输出)已被广泛用于指导基准测试与微调。本文基于Chatbot Arena数据及两名AI裁判(GPT-4o和Llama 3 70B)的判断,对比不同拒绝类型在模型与人类用户间的评分差异。我们区分了伦理拒答(如“我不能协助,因可能造成伤害”)与技术拒答(如“我无法回答,因缺乏实时数据”)。结果表明,模型裁判对伦理拒答的评分显著高于人类用户,而技术拒答则无此差异。我们将这种差异称为‘内容审核偏差’——即模型评估系统系统性地更青睐拒绝行为。这引发了关于透明度、价值对齐及自动化评估中隐含规范假设的深层问题。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly deployed in high-stakes settings, their ability to refuse ethically sensitive prompts-such as those involving hate speech or illegal activities-has become central to content moderation and responsible AI practices. While refusal responses can be viewed as evidence of ethical alignment and safety-conscious behavior, recent research suggests that users may perceive them negatively. At the same time, automated assessments of model outputs are playing a growing role in both evaluation and training. In particular, LLM-as-a-Judge frameworks-in which one model is used to evaluate the output of another-are now widely adopted to guide benchmarking and fine-tuning. This paper examines whether such model-based evaluators assess refusal responses differently than human users. Drawing on data from Chatbot Arena and judgments from two AI judges (GPT-4o and Llama 3 70B), we compare how different types of refusals are rated. We distinguish ethical refusals, which explicitly cite safety or normative concerns (e.g., "I can't help with that because it may be harmful"), and technical refusals, which reflect system limitations (e.g., "I can't answer because I lack real-time data"). We find that LLM-as-a-Judge systems evaluate ethical refusals significantly more favorably than human users, a divergence not observed for technical refusals. We refer to this divergence as a moderation bias-a systematic tendency for model-based evaluators to reward refusal behaviors more than human users do. This raises broader questions about transparency, value alignment, and the normative assumptions embedded in automated evaluation systems.

AI伦理内容审核模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。