arXiv:2502.11246cs.IRcs.CL2025-02中稿 · Transactions on Ma…被引 5

用常识理解图文隐喻,精准识别搞笑背后的歧视与恶意

MemeSense: An Adaptive In-Context Framework for Social Commonsense Driven Meme Moderation

  • 结合视觉文本与常识线索,动态生成适配语境的干预策略
  • 非文字类恶搞图识别准确率提升9%,语义相似度高出35%
  • 适合需要理解文化隐喻的平台内容审核场景

在线表情包是内容审核中极具挑战性的媒介,常以幽默、反讽或文化符号掩盖有害意图。传统依赖显式文本的审核系统难以识别此类隐蔽伤害。我们提出MemeSense,一种自适应框架,通过融合视觉与文本理解,并引入富含常识线索的精选示例,实现对隐含威胁(如性别歧视、刻板印象、低俗内容)的精准检测,即使在无明显文字的图片中亦有效。在多个基准数据集上,MemeSense显著优于现有方法,对非文字类表情包实现最高35%的语义相似度提升和9%的BERTScore改善,同时在文本密集型表情包上也取得显著进步。该研究为构建更安全、更具备上下文感知能力的真实世界内容审核AI系统提供了新路径。代码与数据见:https://github.com/sayantan11995/MemeSense

原文摘要 · Abstract (English)

Online memes are a powerful yet challenging medium for content moderation, often masking harmful intent behind humor, irony, or cultural symbolism. Conventional moderation systems "especially those relying on explicit text" frequently fail to recognize such subtle or implicit harm. We introduce MemeSense, an adaptive framework designed to generate socially grounded interventions for harmful memes by combining visual and textual understanding with curated, semantically aligned examples enriched with commonsense cues. This enables the model to detect nuanced complexed threats like misogyny, stereotyping, or vulgarity "even in memes lacking overt language". Across multiple benchmark datasets, MemeSense outperforms state-of-the-art methods, achieving up to 35% higher semantic similarity and 9% improvement in BERTScore for non-textual memes, and notable gains for text-rich memes as well. These results highlight MemeSense as a promising step toward safer, more context-aware AI systems for real-world content moderation. Code and data available at: https://github.com/sayantan11995/MemeSense

内容审核图像理解常识推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。