arXiv:2506.18919cs.CLcs.AI2025-06被引 8

构建首个带思维链标注的大规模有害梗图数据集,提升模型识别隐含风险能力。

From Recognition to Reasoning: Advancing Multimodal Harmful Meme Detection via Chain-of-Thought Alignment

  • 用思维链标注增强多模态理解,逐步训练模型推理能力。
  • 在自建数据集上准确率超越现有方法,显著提升对隐性危害的识别效果。
  • 适合关注网络内容安全、可解释AI的科研与工程人员使用。

作为融合图像与文本的多模态传播媒介,梗图常通过隐喻、讽刺和幽默传递隐性有害内容,使得有害梗图检测成为一项复杂挑战。尽管近期研究在检测精度和模型可解释性方面取得进展,但大规模高质量的有害梗图数据集仍稀缺,且现有方法在识别隐性风险与细粒度语义理解方面存在明显局限。为此,我们构建了MemeMind,一个面向有害梗图检测的大规模数据集。该数据集包含广泛公开的梗图,并采用依据国际公认标准与当代网络语境制定的全面有害内容分类体系。此外,数据集提供详细的结构化思维链(Chain-of-Thought, CoT)标注,支持对有害性、隐含意图及深层语义的细粒度分析。基于MemeMind,我们进一步提出MemeGuard,一种面向推理的多模态有害梗图检测框架。MemeGuard采用三阶段训练策略,逐步提升模型的视觉理解、多模态推理与有害内容判别能力,从而在检测精度与决策可解释性上均取得提升。大量实验证明,MemeGuard在MemeMind数据集上优于现有最先进方法,为未来有害梗图检测与多模态内容安全研究奠定了坚实基础。

原文摘要 · Abstract (English)

As a multimodal communication medium that integrates images and text, memes often convey implicit harmful content through metaphors, satire, and humor, making harmful meme detection a complex and challenging task. Although recent studies have achieved considerable progress in detection accuracy and model interpretability, large-scale, high-quality datasets for harmful memes remain scarce. Moreover, existing methods still exhibit notable limitations in identifying implicit risks and understanding fine-grained semantics. To address these challenges, we construct MemeMind, a large-scale dataset for harmful meme detection. MemeMind comprises a broad collection of publicly available memes and adopts a rigorous and comprehensive taxonomy of harmful content developed in accordance with widely recognized international standards and contemporary Internet contexts. In addition, the dataset provides detailed structured Chain-of-Thought (CoT) reasoning annotations to support fine-grained analysis of harmfulness, implicit intentions, and underlying semantics in memes. Building upon MemeMind, we further propose MemeGuard, a reasoning-oriented multimodal framework for harmful meme detection. MemeGuard employs a three-stage training strategy to progressively enhance the model's visual understanding, multimodal reasoning, and harmful content discrimination capabilities, thereby improving both detection accuracy and the interpretability of model decisions. Extensive experimental results demonstrate that MemeGuard outperforms existing state-of-the-art methods on the MemeMind dataset, providing a solid foundation for future research on harmful meme detection and multimodal content safety.

有害内容检测多模态思维链数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。