arXiv:2512.17738cs.CL2025-12中稿 · EAMT 2026

UGU翻译评估需明确定义标准程度,否则模型表现不可靠。

When the Gold Standard Isn't Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content

  • 提出12类非标准语言现象与5种翻译策略
  • 不同数据集对UGC处理差异大,影响评分结果
  • 模型表现依赖是否遵循明确翻译指南

用户生成内容(UGC)常含拼写错误、俚语、重复字符和表情符号等非标准表达,导致翻译评估困难:何为‘好’翻译取决于输出的标准程度。我们分析了四个UGC数据集的人工翻译指南,归纳出十二类非标准现象及五种翻译行为(NORMALISE、COPY、TRANSFER、OMIT、CENSOR)。分析显示,各数据集对UGC的处理存在显著差异,参考译文的标准度呈现连续谱。我们发现大型语言模型的翻译评分高度依赖带有明确UGC翻译指令的提示,并在与数据集指南一致时表现更优。因此,公平评估需让模型与评测指标均知晓翻译指南。最后呼吁在数据集构建时制定清晰准则,并开发可控制、指南感知的评估框架。

原文摘要 · Abstract (English)

User-generated content (UGC) is characterised by frequent use of non-standard language, from spelling errors to expressive choices such as slang, character repetitions, and emojis. This makes evaluating UGC translation challenging: what counts as a "good" translation depends on the desired standardness level of the output. To explore this, we examine the human translation guidelines of four UGC datasets, and derive a taxonomy of twelve non-standard phenomena and five translation actions (NORMALISE, COPY, TRANSFER, OMIT, CENSOR). Our analysis reveals notable differences in how UGC is treated, resulting in a spectrum of standardness in reference translations. We show that translation scores of large language models are highly sensitive to prompts with explicit UGC translation instructions, and that they improve when they align with the dataset guidelines. We argue that fair evaluation requires both models and metrics to be aware of translation guidelines. Finally, we call for clear guidelines during dataset creation and for the development of controllable, guideline-aware evaluation frameworks for UGC translation.

翻译评估UGC语言标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。