对比多种机器翻译模型在英-印语对上的表现,评估其通用与专业场景效果。
Evaluating Machine Translation Models for English-Hindi Language Pairs: A Comparative Analysis
- 使用18000+平行语料库和政府问答数据集进行多维度评估
- 不同模型在词汇与学习型指标上表现差异显著
- 适合关注低资源语言翻译性能的研究者或开发者
机器翻译已成为弥合语言差距的关键工具,尤其在英语与印地语这类差异较大的语言之间。本文全面评估了多种用于英-印语互译的机器翻译模型。通过一系列词汇型与基于机器学习的自动评价指标,结合一个包含18000+句对的平行语料库及一个来自政府网站的定制问答数据集,系统比较了不同模型在通用与专业领域中的表现。研究旨在揭示各类翻译方法在处理多样化语言任务时的有效性。结果显示,各模型在不同评价指标下表现不一,凸显了当前系统的优势与改进空间。
原文摘要 · Abstract (English)
Machine translation has become a critical tool in bridging linguistic gaps, especially between languages as diverse as English and Hindi. This paper comprehensively evaluates various machine translation models for translating between English and Hindi. We assess the performance of these models using a diverse set of automatic evaluation metrics, both lexical and machine learning-based metrics. Our evaluation leverages an 18000+ corpus of English Hindi parallel dataset and a custom FAQ dataset comprising questions from government websites. The study aims to provide insights into the effectiveness of different machine translation approaches in handling both general and specialized language domains. Results indicate varying performance levels across different metrics, highlighting strengths and areas for improvement in current translation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。