首个英希机器翻译质量评估基准,助力低资源语言研究。
MTQE.en-he: Machine Translation Quality Estimation for English-Hebrew
- 构建首个公开英希翻译质量评估数据集,含959段双语对与专家评分。
- 三模型集成比最佳单模型提升6.4个百分点(皮尔逊)和5.6个百分点(斯皮尔曼)。
- 参数高效微调方法稳定有效,提升2-3个百分点,适合低资源语言场景。
我们发布MTQE.en-he:据我们所知,首个公开可用的英希机器翻译质量评估基准。该数据集包含来自WMT24++的959个英语片段,每个片段配有一条机器翻译成希伯来语的译文,以及三位人类专家标注的直接评估分数。我们对ChatGPT提示、TransQuest和CometKiwi进行了基准测试,结果表明,三模型集成优于表现最好的单模型(CometKiwi),在皮尔逊相关系数上提升6.4个百分点,在斯皮尔曼相关系数上提升5.6个百分点。对TransQuest和CometKiwi的微调实验显示,全模型更新易受过拟合和分布坍塌影响,而参数高效方法(如LoRA、BitFit和FTHead,即仅微调分类头)训练稳定,并带来2-3个百分点的性能提升。MTQE.en-he及实验结果为该低资源语言对的未来研究提供了支持。
原文摘要 · Abstract (English)
We release MTQE.en-he: to our knowledge, the first publicly available English-Hebrew benchmark for Machine Translation Quality Estimation. MTQE.en-he contains 959 English segments from WMT24++, each paired with a machine translation into Hebrew, and Direct Assessment scores of the translation quality annotated by three human experts. We benchmark ChatGPT prompting, TransQuest, and CometKiwi and show that ensembling the three models outperforms the best single model (CometKiwi) by 6.4 percentage points Pearson and 5.6 percentage points Spearman. Fine-tuning experiments with TransQuest and CometKiwi reveal that full-model updates are sensitive to overfitting and distribution collapse, yet parameter-efficient methods (LoRA, BitFit, and FTHead, i.e., fine-tuning only the classification head) train stably and yield improvements of 2-3 percentage points. MTQE.en-he and our experimental results enable future research on this under-resourced language pair.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。