构建首个孟加拉语网络迷因多模态解释数据集,揭示当前模型在文化隐喻理解上的短板。
BanglaMemeX: Advancing Cultural Metaphoric Image Interpretation in Bangla with a Multimodal Explainable Dataset

- 构建3000条孟加拉语迷因数据集,含幽默、讽刺等多维标注与人工解释
- 现有视觉语言模型在文化隐喻识别上准确率低,表面准确率高但深层理解差
- 适合关注低资源语言、文化敏感多模态系统的研究者使用
视觉语言模型在多模态基准上表现强劲,但对植根于文化的丰富隐喻内容的推理能力仍缺乏深入研究。网络迷因因其图像、叠加文字、反讽和共享社会文化知识之间的隐性互动而构成挑战,而非单纯的视觉识别。这一挑战在低资源语言如孟加拉语中尤为突出,因代码混用、风格化书写及文化特异性符号导致显著分布偏移。本文提出 BanglaMemeX,一个基于文化背景的多模态基准,包含3,000个孟加拉语迷因,附带多维度标签(幽默、讽刺、冒犯性、激励意图、整体情感)以及人类撰写的具体解释,明确描述文本与视觉隐喻。我们系统评估现代视觉语言模型在分类与解释生成任务上的表现,结果表明:尽管表面准确率尚可,当前模型仍难以理解隐性文化线索。研究凸显了具备文化感知能力、能应对语言与文化分布偏移的多模态系统的重要性。
原文摘要 · Abstract (English)
Vision Language Models have achieved strong performance on multimodal benchmarks, yet their ability to reason about culturally grounded and metaphor-rich content remains insufficiently studied. Internet memes present a challenging setting where meaning emerges from implicit interactions between image, overlaid text, sarcasm, and shared socio-cultural knowledge rather than literal visual recognition. This challenge is amplified in low-resource languages such as Bangla, where code-mixing, stylized scripts, and culturally specific symbolism introduce substantial distribution shift. In this work, we introduce BanglaMemeX, a culturally grounded multimodal benchmark comprising 3,000 Bangla memes annotated with multi-dimensional labels (humor, sarcasm, offensiveness, motivational intent, and overall sentiment) and human-written explanations that explicitly describe textual and visual metaphors. We systematically evaluate modern VLMs on both classification and explanation generation, revealing that current models struggle to interpret implicit cultural cues despite reasonable surface-level accuracy. Our results highlight the need for culturally-aware multimodal systems capable of grounded reasoning under linguistic and cultural distribution shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。