用少样本标注数据,让大模型自己推理判断恶搞图片是否有害。
Towards Low-Resource Harmful Meme Detection with LMM Agents
- 让大模型通过检索相似带标签图,借用已有标签信息辅助判断。
- 诱导模型自我修正知识,提炼出泛化性强的有害性判断规律。
- 适合标注数据极少但需精准识别网络恶搞图的场景。
社交媒体中表情包泛滥,亟需有效识别有害内容。由于表情包动态多变,现有数据驱动模型在仅有少量标注样本的低资源场景下表现不佳。本文提出一种基于智能体的低资源有害表情包检测框架,结合外部检索与内部反思策略,利用少量标注样本来增强大模型的多模态推理能力。首先,从数据库中检索相关带标注表情包,将标签信息作为辅助信号输入大模型;随后,激发模型内部的知识修正行为,生成对隐性有害特征的泛化理解。通过整合两种策略,实现对复杂且隐蔽有害模式的辩证推理。在三个表情包数据集上的大量实验表明,该方法在低资源有害表情包检测任务上优于现有最先进方法。
原文摘要 · Abstract (English)
The proliferation of Internet memes in the age of social media necessitates effective identification of harmful ones. Due to the dynamic nature of memes, existing data-driven models may struggle in low-resource scenarios where only a few labeled examples are available. In this paper, we propose an agency-driven framework for low-resource harmful meme detection, employing both outward and inward analysis with few-shot annotated samples. Inspired by the powerful capacity of Large Multimodal Models (LMMs) on multimodal reasoning, we first retrieve relative memes with annotations to leverage label information as auxiliary signals for the LMM agent. Then, we elicit knowledge-revising behavior within the LMM agent to derive well-generalized insights into meme harmfulness. By combining these strategies, our approach enables dialectical reasoning over intricate and implicit harm-indicative patterns. Extensive experiments conducted on three meme datasets demonstrate that our proposed approach achieves superior performance than state-of-the-art methods on the low-resource harmful meme detection task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。