arXiv:2512.03087cs.MMcs.AI2025-12

测试大模型对伪装有害内容的识别能力,发现表现远低于人类。

When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI

  • 构建新基准CamHarmTI,评估图文混合有害内容感知能力。
  • 人类识别准确率超95.75%,主流大模型最高仅2.10%。
  • 微调可提升模型感知力,尤其增强视觉编码器早期层敏感性。

大型视觉语言模型(LVLM)在在线内容审核等任务中日益重要,但真实有害内容常通过图像与文字的微妙互动(如梗图或含恶意文本的图片)伪装隐藏,逃避检测。本文提出CamHarmTI基准,用于评估LVLM对图文组合中伪装有害内容的感知与理解能力,涵盖超过4,500个样本,分三类图像-文本帖子。对100名人类用户和12种主流LVLM的实验显示:人类识别准确率超95.75%,而当前模型普遍失效,例如ChatGPT-4o仅达2.10%。微调实验表明,该基准可有效提升模型感知力,使Qwen2.5VL-7B准确率提升55.94%。注意力分析与逐层探测揭示,微调主要增强视觉编码器早期层的敏感性,促进更整合的场景理解。研究揭示了LVLM固有的感知局限,并为构建更贴近人类的视觉推理系统提供方向。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) are increasingly used for tasks where detecting multimodal harmful content is crucial, such as online content moderation. However, real-world harmful content is often camouflaged, relying on nuanced text-image interplay, such as memes or images with embedded malicious text, to evade detection. This raises a key question: \textbf{can LVLMs perceive such camouflaged harmful content as sensitively as humans do?} In this paper, we introduce CamHarmTI, a benchmark for evaluating LVLM ability to perceive and interpret camouflaged harmful content within text-image compositions. CamHarmTI consists of over 4,500 samples across three types of image-text posts. Experiments on 100 human users and 12 mainstream LVLMs reveal a clear perceptual gap: humans easily recognize such content (e.g., over 95.75\% accuracy), whereas current LVLMs often fail (e.g., ChatGPT-4o achieves only 2.10\% accuracy). Moreover, fine-tuning experiments demonstrate that \bench serves as an effective resource for improving model perception, increasing accuracy by 55.94\% for Qwen2.5VL-7B. Attention analysis and layer-wise probing further reveal that fine-tuning enhances sensitivity primarily in the early layers of the vision encoder, promoting a more integrated scene understanding. These findings highlight the inherent perceptual limitations in LVLMs and offer insight into more human-aligned visual reasoning systems.

多模态安全视觉语言模型内容检测感知缺陷

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。