构建首个针对网络迷因毒性的语义理解评测基准,评估模型对隐含意图和共通知识的推理能力。
M-QUEST -- Meme Question-Understanding Evaluation on Semantics and Toxicity
- 提出包含10个维度的迷因理解框架,涵盖文本、视觉、背景知识与毒性评估等要素
- 构建609个问答对的M-QUEST基准,覆盖307个真实迷因,聚焦共通知识与毒性成因
- 发现指令微调模型在多数维度表现更好,但对语用推断仍存挑战,适合内容安全研究者
网络迷因是重要的在线传播形式,但其依赖常识且多模态的特性使得毒性检测困难。本文提出一个语义理解框架及对应的评测基准,以支持自动提取迷因中的关键意义元素。首先,识别出理解迷因所需的10个维度:文本、视觉、场景、背景知识、情绪、符号投射、类比映射、整体意图、目标群体和毒性评估。其次,基于该框架设计半自动流程,生成包含609个问答对的M-QUEST基准,覆盖307个真实迷因,问题聚焦于毒性判断及其深层原因。最后,评估8个开源大语言模型在该任务上的表现,结果显示当前模型在不同维度和架构间表现差异显著;经指令微调和具备推理能力的模型明显更优,但语用推理类问题依然具有挑战性。代码、数据集与提示已公开,支持多模态内容安全与常识推理方向的研究。
原文摘要 · Abstract (English)
Internet memes are a powerful form of online communication, yet their nature and reliance on commonsense knowledge make toxicity detection challenging. Identifying key features for meme interpretation and understanding, is a crucial task. Previous work has been focused on some elements contributing to the meaning, such as the Textual dimension via OCR, the Visual dimension via object recognition, upper layers of meaning like the Emotional dimension, Toxicity detection via proxy variables, such as hate speech detection, and sentiment analysis. Nevertheless, there is still a lack of an overall architecture able to formally identify elements contributing to the meaning of a meme, and be used in the sense-making process. In this work, we present a semantic framework and a corresponding benchmark for automatic knowledge extraction from memes. First, we identify the necessary dimensions to understand and interpret a meme: Textual material, Visual material, Scene, Background Knowledge, Emotion, Semiotic Projection, Analogical Mapping, Overall Intent, Target Community, and Toxicity Assessment. Second, the framework guides a semi-automatic process of generating a benchmark with commonsense question-answer pairs about meme toxicity assessment and its underlying reason. The resulting benchmark M-QUEST consists of 609 question-answer pairs for 307 memes. Thirdly, we evaluate eight open-source large language models on their ability to correctly solve M-QUEST. Our results show that current models' commonsense reasoning capabilities for toxic meme interpretation vary depending on the dimension and architecture. Models with instruction tuning and reasoning capabilities significantly outperform the others, though pragmatic inference questions remain challenging. We release code, benchmark, and prompts to support future research intersecting multimodal content safety and commonsense reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。