用视觉语言模型实现跨文化梗图再创作,提升幽默与意图传递效果。
Beyond Translation: Cross-Cultural Meme Transcreation with Vision-Language Models
- 融合视觉语言模型的混合生成框架,支持多模态跨文化适配。
- 6315组跨文化梗图对比显示美→中转换质量优于中→美。
- 揭示幽默与图文设计在跨文化中的可迁移性与难点,适合多模态生成研究者。
梗图是网络交流的普遍形式,但其文化特异性给跨文化适应带来挑战。本文研究跨文化梗图再创作,即在保留传播意图和幽默感的同时,适配文化特定引用。提出基于视觉语言模型的混合再创作框架,并构建大规模双向中-美梗图数据集。通过人工评估与自动化指标,分析6315组梗图对在不同文化方向上的再创作质量。结果表明,当前视觉语言模型可在有限范围内实现跨文化再创作,但存在明显方向不对称:美→中转换质量始终高于中→美。进一步识别出幽默元素及图文设计中可迁移的部分,以及仍具挑战性的方面,并提出评估跨文化多模态生成的质量框架。代码与数据集已公开于https://github.com/AIM-SCU/MemeXGen。
原文摘要 · Abstract (English)
Memes are a pervasive form of online communication, yet their cultural specificity poses significant challenges for cross-cultural adaptation. We study cross-cultural meme transcreation, a multimodal generation task that aims to preserve communicative intent and humor while adapting culture-specific references. We propose a hybrid transcreation framework based on vision-language models and introduce a large-scale bidirectional dataset of Chinese and US memes. Using both human judgments and automated evaluation, we analyze 6,315 meme pairs and assess transcreation quality across cultural directions. Our results show that current vision-language models can perform cross-cultural meme transcreation to a limited extent, but exhibit clear directional asymmetries: US-Chinese transcreation consistently achieves higher quality than Chinese-US. We further identify which aspects of humor and visual-textual design transfer across cultures and which remain challenging, and propose an evaluation framework for assessing cross-cultural multimodal generation. Our code and dataset are publicly available at https://github.com/AIM-SCU/MemeXGen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。