提出首个可量化翻译风格的评分指标,支持跨领域通用评估。
Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese
- 用两个微调语言模型的似然比计算翻译风格指数(T-index)。
- 仅需1-5千条合成数据微调,0.5B参数模型即可有效捕捉翻译特征。
- 与主流翻译质量评估指标相关性低,可作为互补测量工具。
Translationese 指通常出现在翻译文本中的语言特征。以往研究将 Translationese 视为原文与译文之间的二元分类问题。本文主张 Translationese 应为连续度量而非二值判断,并提出首个 Translationese 量化指标——翻译风格指数(T-index),通过两个对比微调语言模型(LMs)的似然比计算得出。我们使用合成翻译和真实场景翻译评估 T-index 在跨领域设置下的泛化能力及其与人类判断的一致性。结果表明,T-index 能有效泛化至未见的体裁、作者及语言对。此外,仅用 1-5 千对合成数据微调的两个 0.5B 参数语言模型所计算的 T-index 即可有效捕捉 Translationese 特征,其与人工点评和成对判断高度一致。同时,T-index 与现有机器翻译质量估计(QE)指标如 BLEU、COMET 的相关性较低,说明其不被这些指标覆盖,可在 MT QE 中发挥补充作用。
原文摘要 · Abstract (English)
Translationese refers to linguistic properties that usually occur in translated texts. Previous works study translationese by framing it as a binary classification between original texts and translated texts. In this paper, we argue that translationese should be graded instead of binary and propose the first measure for translationese -- the translationese-index (T-index), computed from the likelihood ratios of two contrastively fine-tuned language models (LMs). We use synthesized translations and translations in the wild to evaluate T-index's generalizability in cross-domain settings and its validity against human judgments. Our results show that T-index can generalize to unseen genres, authors, and language pairs. Moreover, T-index computed using two 0.5B LMs fine-tuned on only 1-5k pairs of synthetic data can effectively capture translationese, as demonstrated by alignment with human pointwise ratings and pairwise judgments. Additionally, the correlation between T-index and existing machine translation (MT) quality estimation (QE) metrics such as BLEU and COMET is low, suggesting that T-index is not covered by these metrics and can serve as a complementary metric in MT QE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。