突破模板限制,提升非模板表情包匹配准确率
Beyond Meme Templates: Limitations of Visual Similarity Measures in Meme Matching
- 提出超越模板的广义表情包匹配方法,覆盖非模板类内容
- 分段相似度计算在非模板表情包上表现优于整体图像匹配
- 基于大模型提示的方法为复杂匹配提供新思路,适合研究数字文化者
网络表情包已成为数字交流的核心形式,在在线社区互动与数字文化研究中具有重要意义。其典型特征是重复使用视觉元素,而基于共享视觉元素的表情包匹配(Meme Matching)是众多分析方法的基础。然而,现有方法多假设每个表情包都有一个固定背景模板(Template)并叠加文字,仅比较背景图,导致无法处理非模板类表情包,限制了自动化分析效果,并难以关联到主流网络表情包词典。本文提出更广泛的匹配范式,突破模板限制。实验表明,传统相似度度量(包括新型分段计算方式)在模板类表情包上表现优异,但在非模板类上显著下降;而分段方法在非模板场景下始终优于整体图像匹配。此外,我们探索了基于预训练多模态大模型的提示式匹配方法。结果表明,仅依赖背景模板的匹配仍存在挑战,需更精细的视觉匹配技术。
原文摘要 · Abstract (English)
Internet memes, now a staple of digital communication, play a pivotal role in how users engage within online communities and allow researchers to gain insight into contemporary digital culture. These engaging user-generated content are characterised by their reuse of visual elements also found in other memes. Matching instances of memes via these shared visual elements, called Meme Matching, is the basis of a wealth of meme analysis approaches. However, most existing methods assume that every meme consists of a shared visual background, called a Template, with some overlaid text, thereby limiting meme matching to comparing the background image alone. Current approaches exclude the many memes that are not template-based and limit the effectiveness of automated meme analysis and would not be effective at linking memes to contemporary web-based meme dictionaries. In this work, we introduce a broader formulation of meme matching that extends beyond template matching. We show that conventional similarity measures, including a novel segment-wise computation of the similarity measures, excel at matching template-based memes but fall short when applied to non-template-based meme formats. However, the segment-wise approach was found to consistently outperform the whole-image measures on matching non-template-based memes. Finally, we explore a prompting-based approach using a pretrained Multimodal Large Language Model for meme matching. Our results highlight that accurately matching memes via shared visual elements, not just background templates, remains an open challenge that requires more sophisticated matching techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。