arXiv:2509.21577cs.CL2025-09被引 1

测试大模型翻译习语时的文化适配度,发现语法对了但意思常错。

"Be My Cheese?": Assessing Cultural Nuance in Multilingual LLM Translations

  • 用87封电商邮件测试24种方言,评估文化贴合度
  • 高资源语言模型仍常误译双关和成语,需人工修正
  • 提出文化恰当性是评估多语言模型的新关键指标

本研究探索前沿多语言AI模型在翻译英语习语、双关等隐喻语言时的本地化能力。针对现有研究过度关注语法正确性和词级准确率的问题,本工作聚焦文化恰当性与整体本地化质量,这对营销和电商等实际应用至关重要。研究选取87个来自20种语言的24种地区方言的电商营销邮件样本,由精通目标语言的人类评审员对翻译结果在语气、含义和受众契合度上的忠实度进行量化评分与定性反馈。结果显示,尽管主流模型普遍生成语法正确的翻译,但对文化隐喻的处理仍存在明显不足,常需大量人工干预。值得注意的是,即使在高资源全球语言中,模型也频繁误译修辞表达。该研究挑战了数据量决定翻译质量的假设,首次将文化恰当性作为衡量多语言大模型性能的核心维度之一,填补了当前学术与产业基准的空白。本初步研究揭示了现有系统在真实本地化场景中的局限,支持未来开展更大规模研究以获得可推广结论,并指导跨文化语境下可靠机器翻译流程的部署。

原文摘要 · Abstract (English)

This pilot study explores the localisation capabilities of state-of-the-art multilingual AI models when translating figurative language, such as idioms and puns, from English into a diverse range of global languages. It expands on existing LLM translation research and industry benchmarks, which emphasise grammatical accuracy and token-level correctness, by focusing on cultural appropriateness and overall localisation quality - critical factors for real-world applications like marketing and e-commerce. To investigate these challenges, this project evaluated a sample of 87 LLM-generated translations of e-commerce marketing emails across 24 regional dialects of 20 languages. Human reviewers fluent in each target language provided quantitative ratings and qualitative feedback on faithfulness to the original's tone, meaning, and intended audience. Findings suggest that, while leading models generally produce grammatically correct translations, culturally nuanced language remains a clear area for improvement, often requiring substantial human refinement. Notably, even high-resource global languages, despite topping industry benchmark leaderboards, frequently mistranslated figurative expressions and wordplay. This work challenges the assumption that data volume is the most reliable predictor of machine translation quality and introduces cultural appropriateness as a key determinant of multilingual LLM performance - an area currently underexplored in existing academic and industry benchmarks. As a proof of concept, this pilot highlights limitations of current multilingual AI systems for real-world localisation use cases. Results of this pilot support the opportunity for expanded research at greater scale to deliver generalisable insights and inform deployment of reliable machine translation workflows in culturally diverse contexts.

多语言模型文化适配本地化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。