arXiv:2509.13608cs.LG2025-09

GPT-4o mini因安全过滤器失效,导致图文仇恨内容检测能力下降。

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection

  • 模型在图文识别中被盲区过滤器打断,失去多模态推理能力。
  • 144次拒绝中,视觉与文本各占50%,说明过滤机制不区分内容风险。
  • 安全系统过于僵化,误伤正常表情包,适合关注AI对齐的研究者。

随着大型多模态模型(LMMs)融入日常数字生活,理解其安全架构成为人工智能对齐的关键问题。本文对全球部署的OpenAI GPT-4o mini进行系统分析,聚焦于复杂的多模态仇恨言论检测任务。基于Hateful Memes Challenge数据集,我们对500个样本展开多阶段调查,探究模型推理与失败模式。核心发现为‘单模态瓶颈’——模型先进多模态推理能力被无上下文感知的安全过滤器系统性中断。144次内容政策拒绝的定量验证显示,视觉与文本内容触发比例均为50%。进一步表明该安全系统脆弱,不仅拦截高风险图像,还误拦常见良性表情包格式,引发可预测的误报。这些结果揭示了当前顶尖LMMs在能力与安全间的根本矛盾,强调需采用更集成、上下文感知的对齐策略,以实现安全且有效的部署。

原文摘要 · Abstract (English)

As Large Multimodal Models (LMMs) become integral to daily digital life, understanding their safety architectures is a critical problem for AI Alignment. This paper presents a systematic analysis of OpenAI's GPT-4o mini, a globally deployed model, on the difficult task of multimodal hate speech detection. Using the Hateful Memes Challenge dataset, we conduct a multi-phase investigation on 500 samples to probe the model's reasoning and failure modes. Our central finding is the experimental identification of a "Unimodal Bottleneck," an architectural flaw where the model's advanced multimodal reasoning is systematically preempted by context-blind safety filters. A quantitative validation of 144 content policy refusals reveals that these overrides are triggered in equal measure by unimodal visual 50% and textual 50% content. We further demonstrate that this safety system is brittle, blocking not only high-risk imagery but also benign, common meme formats, leading to predictable false positives. These findings expose a fundamental tension between capability and safety in state-of-the-art LMMs, highlighting the need for more integrated, context-aware alignment strategies to ensure AI systems can be deployed both safely and effectively.

多模态安全过滤仇恨言论模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。