arXiv:2601.12539cs.AIcs.CL2026-01被引 4

统一多语言多任务的 memes 理解模型,提升跨领域泛化能力。

MemeLens: Multilingual Multitask VLMs for Memes

  • 构建统一标签体系,整合38个数据集的20类任务
  • 实验证明联合训练比单任务微调更鲁棒
  • 适合研究跨文化、多模态内容理解的学者

表情包是网络沟通与操纵的主要载体,其意义源于文本、图像和文化语境的交互。现有研究分散于不同任务(仇恨、性别歧视、宣传、情感、幽默)和语言,限制了跨领域泛化。为此,我们提出 MemeLens,一种统一的多语言多任务解释增强型视觉语言模型(VLM),用于表情包理解。我们整合了38个公开的表情包数据集,过滤并映射数据集特有标签至20个任务的共享分类体系,涵盖危害、目标、隐喻/语用意图及情绪。我们进行了全面的实证分析,涵盖建模范式、任务类别和数据集。结果表明,稳健的表情包理解需要多模态训练,在语义类别间存在显著差异,且模型在仅针对单一数据集微调时易出现过拟合。我们已向社区公开实验资源(https://github.com/MohamedBayan/MemeLens)、模型(https://huggingface.co/QCRI/MemeLens-VLM)和数据集(https://huggingface.co/datasets/QCRI/MemeLens)。

原文摘要 · Abstract (English)

Memes are a dominant medium for online communication and manipulation because meaning emerges from interactions between embedded text, imagery, and cultural context. Existing meme research is distributed across tasks (hate, misogyny, propaganda, sentiment, humour) and languages, which limits cross-domain generalization. To address this gap we propose MemeLens, a unified multilingual and multitask explanation-enhanced Vision Language Model (VLM) for meme understanding. We consolidate $38$ public meme datasets, filter and map dataset-specific labels into a shared taxonomy of $20$ tasks spanning harm, targets, figurative/pragmatic intent, and affect. We present a comprehensive empirical analysis across modeling paradigms, task categories, and datasets. Our findings suggest that robust meme understanding requires multimodal training, exhibits substantial variation across semantic categories, and remains sensitive to over-specialization when models are fine-tuned on individual datasets rather than trained in a unified setting. We make the experimental resources (https://github.com/MohamedBayan/MemeLens), model (https://huggingface.co/QCRI/MemeLens-VLM) and datasets (https://huggingface.co/datasets/QCRI/MemeLens) publicly available to the community.

多模态表情包跨语言VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。