arXiv:2505.23914cs.CLcs.AI2025-05被引 3

发现大模型误判背后隐藏的隐性话题偏见,超越关键词触发。

Probing Association Biases in LLM Moderation Over-Sensitivity

  • 通过上下文诱导测试,量化话题与毒性之间的异常关联模式。
  • 先进模型如GPT-4 Turbo在误判中展现更强话题偏见,尽管整体误判率低。
  • 改变话题提示可显著影响误判率,说明话题框架影响判断决策。

大语言模型广泛用于内容审核,但常出现过度敏感,导致良性内容被误判或安全指令被拒绝。以往研究多归因于显式攻击性词汇,我们通过统计分析发现更深层问题:当对去背景化语句过度敏感时,模型会表现出系统性的主题-毒性关联模式,超越字面触发词。为此提出基于行为的探测方法“主题关联分析”,通过生成简短上下文场景,量化原始评论与场景间的话题放大效应。在多个大模型和大规模数据上验证,更先进的模型(如GPT-4 Turbo)在误报案例中表现出更强的主题关联偏差,尽管其整体误报率更低。通过控制前缀干预实验,证实主题线索可显著改变误报率,表明主题框架是决策相关因素。结果表明,缓解过度敏感需关注模型习得的主题关联,而不仅是关键词过滤。

原文摘要 · Abstract (English)

Large Language Models are widely used for content moderation but often present certain over-sensitivity, leading to misclassification of benign content and rejecting safe user commands. While previous research attributes this issue primarily to the presence of explicit offensive triggers, we statistically reveal a deeper connection beyond token level: When behaving over-sensitively, particularly on decontextualized statements, LLMs exhibit systematic topic-toxicity association patterns that go beyond explicit offensive triggers. To characterize these patterns, we propose Topic Association Analysis, a behavior-based probe that elicits short contextual scenarios for benign inputs and quantifies topic amplification between the scenario and the original comment. Across multiple LLMs and large-scale data, we find that more advanced models (e.g., GPT-4 Turbo) show stronger topic-association skew in false-positive cases despite lower overall false-positive rates. Moreover, via controlled prefix interventions, we show that topic cues can measurably shift false-positive rates, indicating that topic framing is decision-relevant. These results suggest that mitigating over-sensitivity may require addressing learned topic associations in addition to keyword-based filtering.

大模型内容审核偏见检测误判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。