arXiv:2512.21598cs.CV2025-12KDD被引 3

无需标注数据,让AI自进化识别复杂有害梗图

From Shallow Humor to Metaphor: Towards Label-Free Harmful Meme Detection via LMM Agent Self-Improvement

  • 用简单梗图生成伪标签,驱动AI自我优化
  • 对比学习让模型从正负样本对中提炼检测线索
  • 在多个数据集上超越有标注方法,适合快速应对新危害内容

在线媒体中有害梗图的泛滥对公共健康与社会稳定构成重大威胁。现有检测方法严重依赖大规模人工标注数据,需大量人力且难以适应有害内容的持续演化。为此,我们提出ALARM,首个基于大模态模型(LMM)智能体自提升的无标签有害梗图检测框架。其核心创新在于利用“浅层”梗图中的表达信息,迭代提升模型识别更复杂、更隐蔽梗图的能力。ALARM引入基于置信度的显式梗图识别机制,从原始数据集中分离出显性梗图并赋予伪标签;同时提出成对学习引导的智能体自提升范式,将显性梗图重组为正负对比对,以优化学习型LMM智能体。该智能体可自主从对比对中提取高层检测线索,从而有效应对复杂挑战性梗图。在三个多样化数据集上的实验表明,ALARM性能优异且对新出现的梗图具有强适应性,甚至超越部分有标注方法。结果凸显了无标签框架在动态网络环境中应对新型有害内容的可扩展性与前景。

原文摘要 · Abstract (English)

The proliferation of harmful memes on online media poses significant risks to public health and stability. Existing detection methods heavily rely on large-scale labeled data for training, which necessitates substantial manual annotation efforts and limits their adaptability to the continually evolving nature of harmful content. To address these challenges, we present ALARM, the first lAbeL-free hARmful Meme detection framework powered by Large Multimodal Model (LMM) agent self-improvement. The core innovation of ALARM lies in exploiting the expressive information from "shallow" memes to iteratively enhance its ability to tackle more complex and subtle ones. ALARM consists of a novel Confidence-based Explicit Meme Identification mechanism that isolates the explicit memes from the original dataset and assigns them pseudo-labels. Besides, a new Pairwise Learning Guided Agent Self-Improvement paradigm is introduced, where the explicit memes are reorganized into contrastive pairs (positive vs. negative) to refine a learner LMM agent. This agent autonomously derives high-level detection cues from these pairs, which in turn empower the agent itself to handle complex and challenging memes effectively. Experiments on three diverse datasets demonstrate the superior performance and strong adaptability of ALARM to newly evolved memes. Notably, our method even outperforms label-driven methods. These results highlight the potential of label-free frameworks as a scalable and promising solution for adapting to novel forms and topics of harmful memes in dynamic online environments.

有害内容检测自提升无监督学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。