跨文化表情包重创框架,让梗图在不同语言间既有趣又准确。
TransMeme: A Multi-Agent Framework for Cross-Cultural Meme Transcreation

- 用多个专业代理协同完成文化适配、文本重写与图像调整。
- 人工评估提升33.1%,大模型评分胜出率60%领先第二名。
- 适合做跨文化传播、多模态内容生成的研究与实践者。
网络表情包是普遍的多模态在线交流形式,但用户常来自不同语言和文化背景。跨文化表情包重创需同时保留传播意图、适配目标文化的隐含意义,并维持图文一致性。本文首先分析该任务的三大挑战:文化知识理解、意图与语气保留、多模态一致性。基于此,提出多代理框架,包含文化适配、文本重写、修订与条件化视觉调整等专用代理,通过协同反馈处理复杂案例。在中英双向表情包重创上评估,采用人工评价与大模型评分(LLM-as-a-Judge)。结果表明,本方法在两项评价中均优于所有基线。人工评估四项维度全胜,平均提升33.1%;大模型评分中达到60%最高首名率(第二名26%)。各组件均有贡献,错误分析指出当前瓶颈在于幽默重构与图文对齐,而非简单文化知识缺失,提示未来需聚焦幽默迁移研究。
原文摘要 · Abstract (English)
Internet memes are a pervasive form of multimodal online communication; however, such communication often involves users from diverse linguistic and cultural backgrounds. Therefore, adapting memes across cultures and languages is a central challenge for enabling mutual understanding in online communication. Unlike ordinary translation or standalone text rewriting, cross-cultural meme transcreation must jointly preserve communicative intent, adapt culture-dependent meaning for the target audience, and maintain coherence between text and image. In this work, we first provide an explicit task analysis of cross-cultural meme transcreation and identify three core challenges: culture-specific knowledge understanding, intent and tone preservation, and multimodal consistency. Based on this analysis, we propose a multi-agent framework with specialized agents that are coordinated to address these challenges through cultural adaptation, target text rewriting, revision, and conditional visual adjustment. The framework strengthens target text adaptation with coordinated feedback to handle difficult cases that require deeper cultural or visual intervention. We evaluate the framework on bidirectional Chinese-English meme transcreation using both human evaluation and LLM-as-a-Judge. Our method consistently outperforms all baselines across both evaluation settings. In human evaluation, it achieves the best performance on all four dimensions and delivers a 33.1% average improvement over the strongest baseline, while in LLM-as-a-Judge, it attains the highest Top-1 ranking rate (60% versus 26% for the second-best baseline). Further analysis indicates that each component contributes to the performance. Our error analysis suggests that the remaining bottlenecks lie in humor reconstruction and image-text alignment rather than simple cultural knowledge gaps, pointing to future work on humor transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。