首个评估多语言大模型文化翻译能力的真人评测基准
"Be My Cheese?": Cultural Nuance Benchmarking for Machine Translation in Multilingual LLMs
- 构建跨15语言、含文化语境片段的真人评分体系
- 文化概念和节日翻译平均分2.2,俚语和双关仅1.45
- 揭示语法正确但文化失真问题,适合评估本地化模型
我们提出一个大规模真人评估基准,用于衡量先进多语言大模型在机器翻译中的文化本地化能力。现有评测侧重词级与语法准确性,常忽略实际应用所需的语用与文化素养。基于对20种语言87次翻译的试点研究,我们在15种目标语言中评估7个大模型,每语言由5名母语者评分。评分包括全文与分段(习语、双关、节日、文化概念)质量,采用0-3分等级制;分段评分另设'未翻译'选项。全文平均质量为1.68/3:GPT-5(2.10)、Claude Sonnet 4(1.97)、Mistral Medium 3.1(1.84)表现最佳。分段结果显示:节日(2.20/3)与文化概念(2.19/3)优于习语(1.65/3)和双关(1.45/3),后者最易被跳过。使用Krippendorff's α与Gwet's AC2评估,整体一致性中等(α=0.45),双关最低。结果表明语法达标不等于文化契合。该基准为首个聚焦文化细腻度的多语言人工标注评测,强调需更丰富文化数据、改进跨语言语用能力及系统性评估框架。
原文摘要 · Abstract (English)
We present a large-scale human evaluation benchmark for assessing cultural localisation in machine translation produced by state-of-the-art multilingual large language models (LLMs). Existing MT benchmarks emphasise token-level and grammatical accuracy, but often overlook the pragmatic and culturally grounded competencies required for real-world localisation. Building on a pilot study of 87 translations across 20 languages, we evaluate 7 multilingual LLMs across 15 target languages with 5 native-speaker raters per language. Each rater scored both full-text translations and segment-level instances of culturally nuanced language (idioms, puns, holidays, and culturally embedded concepts) on an ordinal 0-3 quality scale; segment ratings additionally included an NA option for untranslated segments. Across full-text evaluations, mean overall quality is modest (1.68/3): GPT-5 (2.10/3), Claude Sonnet 4 (1.97/3), and Mistral Medium 3.1 (1.84/3) form the strongest tier with fewer catastrophic failures. Segment-level results show sharp category effects: holidays (2.20/3) and cultural concepts (2.19/3) translate notably better than idioms (1.65/3) and puns (1.45/3), and idioms are most likely to be left untranslated. Inter-rater reliability was assessed using Krippendorff's α and Gwet's AC2, indicating moderate agreement overall (Krippendorff's α = 0.45) with the lowest agreement for puns. These findings demonstrate a persistent gap between grammatical adequacy and cultural resonance. To our knowledge, this is the first multilingual, human-annotated benchmark focused explicitly on cultural nuance in translation and localisation. The results highlight the need for culturally informed training data, improved cross-lingual pragmatics, and evaluation frameworks that support systematic benchmarking of culturally grounded translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。