用专业术语体系测试ChatGPT自动标注翻译错误的能力
Testing LLMs' Capabilities in Annotating Translations Based on an Error Typology Designed for LSP Translation: First Experiments with ChatGPT
- 基于自定义错误分类体系,用提示词控制评估精度
- 对DeepL译文标注准确率高,但自评表现差
- 适合用于专业翻译教学中的自动化评估辅助
本研究探讨大语言模型(如ChatGPT)在基于专门化错误分类体系下标注机器翻译输出的能力。与以往聚焦通用语言的研究不同,本文关注ChatGPT识别和分类专业领域翻译错误的表现。通过两种不同提示词设计,将ChatGPT的标注结果与人工专家对DeepL和ChatGPT自身生成译文的评估进行对比。结果显示:对于DeepL生成的译文,召回率和精确率均较高;但错误分类的准确性取决于提示词的具体特征和详细程度,使用详细提示时表现优异。而在评估自身生成的译文时,ChatGPT表现显著下降,暴露出自评估的局限性。研究揭示了大语言模型在专业翻译评价中的潜力与不足,为未来基于开源大模型的高质量自动标注提供了方向。后续计划探索该方法在翻译教学中的实际应用,包括优化教师人工评估流程,以及分析模型标注对学生译后编辑与学习的影响。
原文摘要 · Abstract (English)
This study investigates the capabilities of large language models (LLMs), specifically ChatGPT, in annotating MT outputs based on an error typology. In contrast to previous work focusing mainly on general language, we explore ChatGPT's ability to identify and categorise errors in specialised translations. By testing two different prompts and based on a customised error typology, we compare ChatGPT annotations with human expert evaluations of translations produced by DeepL and ChatGPT itself. The results show that, for translations generated by DeepL, recall and precision are quite high. However, the degree of accuracy in error categorisation depends on the prompt's specific features and its level of detail, ChatGPT performing very well with a detailed prompt. When evaluating its own translations, ChatGPT achieves significantly poorer results, revealing limitations with self-assessment. These results highlight both the potential and the limitations of LLMs for translation evaluation, particularly in specialised domains. Our experiments pave the way for future research on open-source LLMs, which could produce annotations of comparable or even higher quality. In the future, we also aim to test the practical effectiveness of this automated evaluation in the context of translation training, particularly by optimising the process of human evaluation by teachers and by exploring the impact of annotations by LLMs on students' post-editing and translation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。