无需标注数据,用多智能体协作识别新型有害表情包。
MIND: A Multi-agent Framework for Zero-shot Harmful Meme Detection
- 从无标注图库中找相似表情包提供上下文。
- 双向推理机制全面理解相似内容特征。
- 多智能体辩论确保判断可靠,适合新类型内容。
社交媒体上表情包的快速传播凸显了检测有害内容的迫切需求。然而,传统数据驱动方法因内容持续演化和缺乏实时标注数据,难以应对新型表情包。为此,我们提出MIND——一种无需标注数据的零样本有害表情包检测多智能体框架。该框架包含三项核心策略:1)从无标注参考集检索相似表情包以提供上下文信息;2)提出双向洞察推导机制,全面理解相似内容;3)采用多智能体辩论机制,通过理性仲裁实现稳健决策。在三个表情包数据集上的大量实验表明,该框架不仅优于现有零样本方法,且在不同模型架构与参数规模下均表现出强泛化能力,为有害表情包检测提供了可扩展解决方案。代码已公开于 https://github.com/destroy-lonely/MIND。
原文摘要 · Abstract (English)
The rapid expansion of memes on social media has highlighted the urgent need for effective approaches to detect harmful content. However, traditional data-driven approaches struggle to detect new memes due to their evolving nature and the lack of up-to-date annotated data. To address this issue, we propose MIND, a multi-agent framework for zero-shot harmful meme detection that does not rely on annotated data. MIND implements three key strategies: 1) We retrieve similar memes from an unannotated reference set to provide contextual information. 2) We propose a bi-directional insight derivation mechanism to extract a comprehensive understanding of similar memes. 3) We then employ a multi-agent debate mechanism to ensure robust decision-making through reasoned arbitration. Extensive experiments on three meme datasets demonstrate that our proposed framework not only outperforms existing zero-shot approaches but also shows strong generalization across different model architectures and parameter scales, providing a scalable solution for harmful meme detection. The code is available at https://github.com/destroy-lonely/MIND.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。