arXiv:2510.05113cs.CLcs.AI2025-10被引 1

为古吉拉特语机器翻译设计可训练的参考评估指标,提升与人工评价的一致性。

Trainable Reference-Based Evaluation Metric for Identifying Quality of English-Gujarati Machine Translation System

  • 基于监督学习构建古吉拉特语专用评估指标,使用25个特征训练。
  • 两个模型分别采用6层和10层隐藏层,均训练500轮,性能优于现有方法。
  • 在7个系统1000条翻译上测试,与人工评分相关性更高,适合印度语言评估。

机器翻译(MT)评估是MT开发流程中的关键环节。缺乏对翻译输出的分析,就无法准确评估系统性能。实验表明,适用于英语等欧洲语言的评估方法在印度语言中效果不佳。本文提出一种基于参考的古吉拉特语机器翻译评估指标,采用监督学习方法。我们训练了两种版本的指标,均使用25个特征,一个模型含6层隐藏层,训练500轮;另一个含10层隐藏层,同样训练500轮。为测试该指标性能,我们收集了7个机器翻译系统的1000条输出,并与1份人工参考译文进行对比。与现有评估指标相比,本指标在人类评分相关性方面表现更优。

原文摘要 · Abstract (English)

Machine Translation (MT) Evaluation is an integral part of the MT development life cycle. Without analyzing the outputs of MT engines, it is impossible to evaluate the performance of an MT system. Through experiments, it has been identified that what works for English and other European languages does not work well with Indian languages. Thus, In this paper, we have introduced a reference-based MT evaluation metric for Gujarati which is based on supervised learning. We have trained two versions of the metric which uses 25 features for training. Among the two models, one model is trained using 6 hidden layers with 500 epochs while the other model is trained using 10 hidden layers with 500 epochs. To test the performance of the metric, we collected 1000 MT outputs of seven MT systems. These MT engine outputs were compared with 1 human reference translation. While comparing the developed metrics with other available metrics, it was found that the metrics produced better human correlations.

机器翻译评估指标古吉拉特语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。