跨语言多模态识别网络迷因中的角色,挑战文化语境与混合语言的精准理解
Decoding Memes: Benchmarking Narrative Role Classification across Multilingual and Multimodal Models
- 构建平衡数据集,对比真实迷因与合成仇恨内容的语言特征
- 大模型如DeBERTa-v3在零样本下表现较好,但'受害者'类识别仍困难
- 结构化提示+角色定义可微调提升多模态模型表现,适合内容安全研究者
本文研究了在互联网迷因中识别叙事角色(英雄、反派、受害者、其他)这一挑战性任务,涵盖英语及英印混合语言的三组测试集。基于原本偏向‘其他’类别的标注数据集,我们采用更均衡且语言多样化的扩展版本,该版本最初作为CLEF 2024共享任务的一部分提出。全面的词汇与结构分析显示,真实迷因使用细腻、文化特异且上下文丰富的语言,而合成仇恨内容则呈现明确且重复的词汇标记。为评估角色检测性能,我们测试了多种模型:微调的多语言Transformer、情感与滥用检测分类器、指令微调的大语言模型及多模态视觉-语言模型,并在零样本设置下以精确率、召回率和F1值进行评估。尽管如DeBERTa-v3和Qwen2.5-VL等大模型表现出显著优势,但对‘受害者’类别的识别仍存在持续困难,跨文化与代码混合内容的泛化能力有限。我们还探索了提示设计策略,发现结合结构化指令与角色定义的混合提示带来边际但一致的改进。结果强调文化背景、提示工程与多模态推理在建模视觉-文本内容微妙叙事框架中的重要性。
原文摘要 · Abstract (English)
This work investigates the challenging task of identifying narrative roles - Hero, Villain, Victim, and Other - in Internet memes, across three diverse test sets spanning English and code-mixed (English-Hindi) languages. Building on an annotated dataset originally skewed toward the 'Other' class, we explore a more balanced and linguistically diverse extension, originally introduced as part of the CLEF 2024 shared task. Comprehensive lexical and structural analyses highlight the nuanced, culture-specific, and context-rich language used in real memes, in contrast to synthetically curated hateful content, which exhibits explicit and repetitive lexical markers. To benchmark the role detection task, we evaluate a wide spectrum of models, including fine-tuned multilingual transformers, sentiment and abuse-aware classifiers, instruction-tuned LLMs, and multimodal vision-language models. Performance is assessed under zero-shot settings using precision, recall, and F1 metrics. While larger models like DeBERTa-v3 and Qwen2.5-VL demonstrate notable gains, results reveal consistent challenges in reliably identifying the 'Victim' class and generalising across cultural and code-mixed content. We also explore prompt design strategies to guide multimodal models and find that hybrid prompts incorporating structured instructions and role definitions offer marginal yet consistent improvements. Our findings underscore the importance of cultural grounding, prompt engineering, and multimodal reasoning in modelling subtle narrative framings in visual-textual content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。