构建爱沙尼亚语常识推理数据集,验证人工与机器翻译对大模型表现的影响。
Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translation
- 用专业译者将WinoGrande译为爱沙尼亚语并适配文化背景。
- 人工翻译集上模型表现略降,机器翻译集表现显著更差。
- 提示工程改善有限,强调语言专家参与翻译的必要性。
本文构建了广泛使用的常识推理基准WinoGrande测试集的爱沙尼亚语本地化版本,由专业译者完成翻译与文化适配。我们评估了专有及开源模型在人工翻译数据集上的表现,并探索将人工翻译经验融入提示设计以提升机器翻译质量的可行性。该提示针对爱沙尼亚语语言特征和WinoGrande特有的翻译挑战进行定制。结果表明,模型在人工翻译的爱沙尼亚语数据集上表现略低于原始英文测试集,而在机器翻译数据集上表现显著更差。实验还显示,提示工程对翻译质量或模型准确率的提升有限,凸显语言专家在数据集翻译与适配中的关键作用,以确保对大语言模型语言能力与推理能力评估的可靠性与可解释性。
原文摘要 · Abstract (English)
In this paper, we present a localized and culturally adapted Estonian translation of the test set from the widely used commonsense reasoning benchmark, WinoGrande. We detail the translation and adaptation process carried out by translation specialists and evaluate the performance of both proprietary and open source models on the human translated benchmark. Additionally, we explore the feasibility of achieving high-quality machine translation by incorporating insights from the manual translation process into the design of a detailed prompt. This prompt is specifically tailored to address both the linguistic characteristics of Estonian and the unique translation challenges posed by the WinoGrande dataset. Our findings show that model performance on the human translated Estonian dataset is slightly lower than on the original English test set, while performance on machine-translated data is notably worse. Additionally, our experiments indicate that prompt engineering offers limited improvement in translation quality or model accuracy, and highlight the importance of involving language specialists in dataset translation and adaptation to ensure reliable and interpretable evaluations of language competency and reasoning in large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。