arXiv:2605.25626cs.CL2026-05中稿 · ICML

评估社交媒体内容翻译的文化适配度,发现大模型更懂文化隐喻。

Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC

论文配图:Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC
图 1 · 摘自论文原文
  • 构建1002条带文化符号的UGC数据集,区分四类表达风格。
  • 测试15个模型发现传统指标无法衡量文化传递效果,模型越大越能懂文化。
  • 开源评测平台与裁判模型,支持实时提交与评分。

社交媒体促进跨语言交流,但用户生成内容(UGC)因非正式语体、文化引用和互动表达,翻译仍具挑战。现有基准与指标常忽略译文是否传达原意及文化共鸣。本文提出CULTURE-MT,一个聚焦文化传递与UGC情感共振的社交媒体翻译评测集,包含1,002条跨14个领域的UGC笔记,按文化负载符号与语言风格分为四类。我们基于UGC特点构建训练数据,微调Qwen3-8B与Qwen3-32B作为基线。提出“文化有效性”新评估标准,关注表达准确性和文化适应性。测试15个模型发现,传统指标无法捕捉文化有效性;且基础大模型的文化有效性与其规模正相关。本研究提供完整评估体系,并开放评测平台与在线排行榜,支持提交结果由训练好的JUDGER模型自动评估。

原文摘要 · Abstract (English)

Social media platforms enable large-scale cross-lingual communication, but translating user-generated content (UGC) remains challenging due to its informal style, cultural references, and interaction-based expressions. While recent LLMs have improved translation quality, existing benchmarks and metrics often fail to capture whether translations convey intended meaning and cultural resonance in real-world settings. In this work, we introduce CULTURE-MT, a benchmark for social media translation that focuses on both CULtural Transmission and UGC-specific emotion REsonance. CULTURE-MT consists of 1,002 UGC notes across 14 domains, categorized into four types based on culture-loaded symbols and linguistic style features. We also construct UGC-oriented training data to fine-tune Qwen3-8B and Qwen3-32B as baselines. We propose cultural effectiveness as a new evaluation criterion, focusing on expression accuracy and cultural adaptability. Testing 15 models, including the baselines, we find that traditional metrics fail to capture cultural effectiveness. We also observe that cultural effectiveness on base LLMs correlates with model size. Our work provides a comprehensive evaluation system for UGC translation models and will offer an open evaluation platform to advance research in this area. We release the CULTURE-MT benchmark and provide an online leaderboard where submitted translation results can be evaluated by our trained JUDGER.

自然语言处理文化翻译大模型评估社交媒体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。