机器翻译会因文化差异导致语义标签偏移,影响下游应用
Semantic Label Drift in Cross-Cultural Translation
- 通过跨文化实验发现翻译中语义标签会因文化差异发生偏移
- 大模型虽具文化知识,但反而加剧了标签漂移现象
- 源语言与目标语言的文化相似性直接影响标签保真度
机器翻译(MT)常被用于低资源语言的合成数据生成。尽管情感保留已受到长期关注,但源语言与目标语言之间的文化对齐问题仍被忽视。本文假设:由于文化差异,语义标签在翻译过程中可能发生漂移。通过在文化敏感与中性领域开展系列实验,我们得出三个关键发现:(1) 包括现代大语言模型(LLMs)在内的MT系统在文化敏感领域易引发标签漂移;(2) 与早期统计模型不同,LLMs具备文化知识,利用这些知识反而会放大标签漂移;(3) 源语言与目标语言之间的文化相似性是标签保真性的关键决定因素。研究揭示,忽略文化因素不仅损害标签准确性,还可能在下游应用中引发误解与文化冲突。
原文摘要 · Abstract (English)
Machine Translation (MT) is widely employed to address resource scarcity in low-resource languages by generating synthetic data from high-resource counterparts. While sentiment preservation in translation has long been studied, a critical but underexplored factor is the role of cultural alignment between source and target languages. In this paper, we hypothesize that semantic labels are drifted or altered during MT due to cultural divergence. Through a series of experiments across culturally sensitive and neutral domains, we establish three key findings: (1) MT systems, including modern Large Language Models (LLMs), induce label drift during translation, particularly in culturally sensitive domains; (2) unlike earlier statistical MT tools, LLMs encode cultural knowledge, and leveraging this knowledge can amplify label drift; and (3) cultural similarity or dissimilarity between source and target languages is a crucial determinant of label preservation. Our findings highlight that neglecting cultural factors in MT not only undermines label fidelity but also risks misinterpretation and cultural conflict in downstream applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。