用谐音词测试翻译系统对中文情感的保留能力
Automatically Generating Chinese Homophone Words to Probe Machine Translation Estimation Systems
- 基于信息论生成挑战性情感谐音词
- 揭示翻译模型在情感保持上的脆弱性
- 适合评估中文UGC翻译质量的模型研究者
评估用户生成内容(UGC)的机器翻译质量面临独特挑战,如源文本情感细微差别是否在目标文本中得以保留。现有研究已提出情感相关数据集、框架和模型,可无需参考译文自动评估中文UGC的翻译质量。然而,这些模型对情感细微差别的鲁棒性尚未充分探索。为此,我们提出一种受信息论启发的新方法,通过自信息概念生成与情感相关的中文谐音词。该方法生成的谐音词曾导致翻译在情感保留上出错,暴露出翻译系统及其评估方法在处理情感化UGC时的缺陷。我们通过人工评估验证生成谐音词的质量,并与现有方法对比,结果显示本方法与人类判断的相关性更高。生成的谐音词及其人工译文被用于构造扰动,以探测现有质量评估模型的鲁棒性,包括多任务学习训练模型、微调的多语言语言模型及大语言模型(LLMs)。结果表明,参数量更大的LLMs在面对此类扰动时表现出更高稳定性与鲁棒性。数据与代码已公开,支持复现与进一步研究。
原文摘要 · Abstract (English)
Evaluating machine translation (MT) of user-generated content (UGC) involves unique challenges such as checking whether the nuance of emotions from the source are preserved in the target text. Recent studies have proposed emotion-related datasets, frameworks and models to automatically evaluate MT quality of Chinese UGC, without relying on reference translations. However, whether these models are robust to the challenge of preserving emotional nuances has been left largely unexplored. To address this gap, we introduce a novel method inspired by information theory which generates challenging Chinese homophone words related to emotions, by leveraging the concept of self-information. Our approach generates homophones that were observed to cause translation errors in emotion preservation, and exposes vulnerabilities in MT systems and their evaluation methods when tackling emotional UGC. We evaluate the efficacy of our method using human evaluation for the quality of these generated homophones, and compare it with an existing one, showing that our method achieves higher correlation with human judgments. The generated Chinese homophones, along with their manual translations, are utilized to generate perturbations and to probe the robustness of existing quality evaluation models, including models trained using multi-task learning, fine-tuned variants of multilingual language models, as well as large language models (LLMs). Our results indicate that LLMs with larger size exhibit higher stability and robustness to such perturbations. We release our data and code for reproducibility and further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。