arXiv:2601.04567cs.CV2026-01ACL

通过复现设计原理,提升对动态变化的有害表情包的检测能力

All Changes May Have Invariant Principles: Improving Ever-Shifting Harmful Meme Detection via Design Concept Reproduction

  • 基于攻击树构建设计概念图,捕捉有害表情包的生成逻辑
  • 在跨类型和时间演进场景下仍保持81.1%准确率,下降轻微
  • 辅助人工发现效率提升,每张表情包节省15~30秒

互联网社区中的有害表情包具有类型更迭和时序演化的特性,难以分析。尽管形式不断变化,我们发现不同表情包可能共享不变的设计原理——即恶意创作者的底层构思。为此,本文提出基于设计概念复现的有害表情包检测方法RepMD。首先,借鉴攻击树构建设计概念图(DCG),描述生成有害表情包的潜在步骤;接着,通过设计步骤复现与图剪枝从历史数据中推导出DCG;最后,利用DCG引导多模态大模型(MLLM)进行检测。实验表明,RepMD在基准测试中达到最高81.1%准确率,在跨类型与时间演化场景下性能下降轻微。人工评估显示,该方法可使人工发现有害表情包的效率提升,单个判断耗时缩短至15~30秒。

原文摘要 · Abstract (English)

Harmful memes are ever-shifting in the Internet communities, which are difficult to analyze due to their type-shifting and temporal-evolving nature. Although these memes are shifting, we find that different memes may share invariant principles, i.e., the underlying design concept of malicious users, which can help us analyze why these memes are harmful. In this paper, we propose RepMD, an ever-shifting harmful meme detection method based on the design concept reproduction. We first refer to the attack tree to define the Design Concept Graph (DCG), which describes steps that people may take to design a harmful meme. Then, we derive the DCG from historical memes with design step reproduction and graph pruning. Finally, we use DCG to guide the Multimodal Large Language Model (MLLM) to detect harmful memes. The evaluation results show that RepMD achieves the highest accuracy with 81.1% and has slight accuracy decreases when generalized to type-shifting and temporal-evolving memes. Human evaluation shows that RepMD can improve the efficiency of human discovery on harmful memes, with 15$\sim$30 seconds per meme.

有害内容检测多模态模型设计原理动态识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。