arXiv:2603.17303cs.CLcs.AI2026-03

构建首个系统性文化表达翻译评测基准,揭示大模型在跨文化语义理解上的深层缺陷。

From Words to Worlds: Benchmarking Cross-Cultural Cultural Understanding in Machine Translation

  • 设计覆盖7959条实例的CulT-Eval基准,涵盖习语、俚语等多类文化表达
  • 发现现有模型在保留文化语义上普遍失败,错误模式系统性存在
  • 提出新评估指标,专门捕捉传统指标忽略的文化意义偏差

文化表达(如习语、俚语、文化特有项)在自然语言中广泛存在,其含义超越字面形式。机器翻译系统准确处理此类表达仍具挑战性。现有评测体系零散,缺乏系统框架来评估文化负载表达的翻译表现。为此,我们提出CulT-Eval,一个用于评估模型处理各类文化表达能力的基准。该基准包含超过7,959个精心构建的实例,覆盖多种文化表达类型,并建立全面的错误分类体系。通过对大语言模型的广泛评测与深入分析,我们识别出若干重复出现且系统性的失败模式,这些模式无法被现有自动评估指标充分捕捉。因此,我们提出一种互补性评估指标,专门针对标准机器翻译指标所忽视的文化引发的意义偏离。结果表明,当前模型难以保持文化相关语义,也未能把握准确翻译所必需的文化与上下文细微差别。我们的基准与代码已公开于https://anonymous.4open.science/r/CulT-Eval-E75D/。

原文摘要 · Abstract (English)

Culture-expressions, such as idioms, slang, and culture-specific items (CSIs), are pervasive in natural language and encode meanings that go beyond literal linguistic form. Accurately translating such expressions remains challenging for machine translation systems. Despite this, existing benchmarks remain fragmented and do not provide a systematic framework for evaluating translation performance on culture-loaded expressions. To address this gap, we introduce CulT-Eval, a benchmark designed to evaluate how models handle different types of culturally grounded expressions. CulT-Eval comprises over 7,959 carefully curated instances spanning multiple types of culturally grounded expressions, with a comprehensive error taxonomy covering culturally grounded expressions. Through extensive evaluation of large language models and detailed analysis, we identify recurring and systematic failure modes that are not adequately captured by existing automatic metrics. Accordingly, we propose a complementary evaluation metric that targets culturally induced meaning deviations overlooked by standard MT metrics. The results indicate that current models struggle to preserve culturally grounded meaning and to capture the cultural and contextual nuances essential for accurate translation. Our benchmark and code are available at https://anonymous.4open.science/r/CulT-Eval-E75D/.

机器翻译文化理解评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。