arXiv:2511.04977cs.CVcs.MM2025-11

构建首个贴纸语义相似度评估基准与通用编码模型。

GSE: Evaluating Sticker Visual Semantic Similarity via a General Sticker Encoder

  • 提出通用贴纸编码器GSE,融合多数据集学习贴纸嵌入。
  • 在905对标注贴纸上表现优于现有视觉模型。
  • 适合研究贴纸理解、检索及多模态生成的学者使用。

贴纸作为流行视觉表达形式,其语义关系因内容高度多样且具象征性而难以理解。本文正式定义贴纸语义相似度任务,提出首个基准Triple-S,包含905对人工标注的正负贴纸对。实验证明,现有预训练视觉与多模态模型难以捕捉贴纸细微语义。为此,我们提出轻量级通用贴纸编码器GSE,利用Triple-S及额外数据集学习鲁棒贴纸嵌入。GSE在未见贴纸上表现更优,并在情绪分类和贴纸检索等下游任务中取得良好效果。通过公开发布Triple-S与GSE,我们提供标准化评估工具与可靠嵌入,推动贴纸理解、检索与多模态内容生成研究。相关资源已公开可获取。

原文摘要 · Abstract (English)

Stickers have become a popular form of visual communication, yet understanding their semantic relationships remains challenging due to their highly diverse and symbolic content. In this work, we formally {define the Sticker Semantic Similarity task} and introduce {Triple-S}, the first benchmark for this task, consisting of 905 human-annotated positive and negative sticker pairs. Through extensive evaluation, we show that existing pretrained vision and multimodal models struggle to capture nuanced sticker semantics. To address this, we propose the {General Sticker Encoder (GSE)}, a lightweight and versatile model that learns robust sticker embeddings using both Triple-S and additional datasets. GSE achieves superior performance on unseen stickers, and demonstrates strong results on downstream tasks such as emotion classification and sticker-to-sticker retrieval. By releasing both Triple-S and GSE, we provide standardized evaluation tools and robust embeddings, enabling future research in sticker understanding, retrieval, and multimodal content generation. The Triple-S benchmark and GSE have been publicly released and are available here.

贴纸理解语义相似度多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。