为英语-希伯来语翻译质量评估构建半合成平行数据集,提升低资源语言对的评估能力。
Semi-Synthetic Parallel Data for Translation Quality Estimation: A Case Study of Dataset Building for an Under-Resourced Language Pair
- 基于典型语言模式生成英文句子,经多引擎翻译并用BLEU筛选,构建半合成数据集。
- 模型在该数据集上表现受数据量、分布均衡性及错误类型分布影响显著。
- 适合研究低资源语言对翻译质量评估、尤其关注形态复杂语言的研究者。
翻译质量评估(QE)在机器翻译流程中至关重要,用于评估无参考译文的输出,并判断是否需要人工润色或重译。然而,由于平行语料有限及语言特性差异(如形态结构复杂的语言),针对低资源语言对的高精度、可适应、可靠的QE系统仍难以实现。本研究构建了用于英-希伯来语QE的半合成平行数据集:基于典型语言模式生成英文句子,通过多个机器翻译引擎翻译为希伯来语,并采用基于BLEU的筛选方法进行过滤;每个译文段落均由语言学家人工评估打分;同时引入我们自有资源中的专业翻译段落,赋予最高质量评分。为应对语言挑战(尤其是性别与数的一致性问题),在译文中引入可控翻译错误。使用BERT和XLM-R等神经网络模型在此数据集上训练,评估句级翻译质量。实验表明,数据集规模、分布均衡性及错误分布对模型性能有显著影响。研究详述了构建挑战、方法与结果,并指明未来改进方向。本工作推动了低资源语言对(包括形态丰富的语言)的QE模型发展。
原文摘要 · Abstract (English)
Quality estimation (QE) plays a crucial role in machine translation (MT) workflows, as it serves to evaluate generated outputs that have no reference translations and to determine whether human post-editing or full retranslation is necessary. Yet, developing highly accurate, adaptable and reliable QE systems for under-resourced language pairs remains largely unsolved, due mainly to limited parallel corpora and to diverse language-dependent factors, such as with morphosyntactically complex languages. This study presents a semi-synthetic parallel dataset for English-to-Hebrew QE, generated by creating English sentences based on examples of usage that illustrate typical linguistic patterns, translating them to Hebrew using multiple MT engines, and filtering outputs via BLEU-based selection. Each translated segment was manually evaluated and scored by a linguist, and we also incorporated professionally translated English-Hebrew segments from our own resources, which were assigned the highest quality score. Controlled translation errors were introduced to address linguistic challenges, particularly regarding gender and number agreement, and we trained neural QE models, including BERT and XLM-R, on this dataset to assess sentence-level MT quality. Our findings highlight the impact of dataset size, distributed balance, and error distribution on model performance. We will describe the challenges, methodology and results of our experiments, and specify future directions aimed at improving QE performance. This research contributes to advancing QE models for under resourced language pairs, including morphology-rich languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。