arXiv:2505.00056cs.CLcs.IR2025-05被引 1

用模板匹配与多维相似性聚类网络迷因,提升准确性与可解释性。

Clustering Internet Memes Through Template Matching and Multi-Dimensional Similarity

  • 基于模板匹配和多维度特征聚类迷因,无需预设数据库。
  • 在形式、视觉、文本、身份等维度上实现更一致的聚类结果。
  • 适合研究迷因传播、毒性检测与内容类型分析的学者使用。

迷因聚类对毒性检测、传播建模和类型识别至关重要,但此前研究关注较少。由于迷因具有多模态性、文化背景依赖性和高度可变性,聚类极具挑战。现有方法依赖预定义数据库,忽略语义信息,难以处理多种相似性维度。本文提出一种新方法:通过模板匹配结合多维相似性特征,无需预先设定数据库,支持自适应匹配。利用形式、视觉内容、文本和身份等多维度的局部与全局特征进行聚类,结果显著优于现有方法,生成的聚类更一致、更连贯。相似性特征集具备可扩展性,符合人类直觉。所有代码已公开,便于后续研究。

原文摘要 · Abstract (English)

Meme clustering is critical for toxicity detection, virality modeling, and typing, but it has received little attention in previous research. Clustering similar Internet memes is challenging due to their multimodality, cultural context, and adaptability. Existing approaches rely on databases, overlook semantics, and struggle to handle diverse dimensions of similarity. This paper introduces a novel method that uses template-based matching with multi-dimensional similarity features, thus eliminating the need for predefined databases and supporting adaptive matching. Memes are clustered using local and global features across similarity categories such as form, visual content, text, and identity. Our combined approach outperforms existing clustering methods, producing more consistent and coherent clusters, while similarity-based feature sets enable adaptability and align with human intuition. We make all supporting code publicly available to support subsequent research.

迷因聚类多模态模板匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。