通过课堂实践,研究学生如何评估AI翻译系统并做出选择。
Evaluative Judgement in Teaching AI-based Translation: A Class-room Case Study of AI-Mediated Translation and Post-Editing
- 让学生对比多个AI翻译系统输出,自主评估并选择最佳结果。
- 学生选择与自动评分排名不符,更关注语义准确、表达自然等质量因素。
- 适合对AI辅助翻译教学感兴趣的教师和研究者参考。
基于某本科翻译课程中23个匿名学生项目,本文研究了在真实课堂任务中,通过对比通用大语言模型与在线机器翻译系统,如何激发学生对AI辅助翻译的评价判断。学生将短篇专业英文维基百科文本翻译为加泰罗尼亚语或西班牙语,生成四种系统输出,使用自动指标及人工准确性/流畅性评估进行打分,选定一个输出进行后期编辑,并在书面报告中说明理由。所有23个项目均报告描述性统计,22个附有书面报告的案例用于质性分析。结果显示,学生并未将自动指标视为最终依据:最终选择常与指标排名不一致,其决策基于语义准确、语言流畅、术语使用、表达自然度及后期编辑工作量等因素。本研究未在受控条件下进行系统基准测试,而是分析学生在真实课堂任务中如何解释其系统选择行为。
原文摘要 · Abstract (English)
Drawing on 23 anonymized student pro-jects from a fourth-year Machine Transla-tion and Post-editing course in a BA-level translation programme, this paper exam-ines how structured comparison of gen-eral-purpose LLMs and online MT sys-tems can elicit evaluative judgement in AI-mediated translation. Students translat-ed short specialised English Wikipedia texts into Catalan or Spanish, generated four system outputs, evaluated them using automatic metrics and human adequa-cy/fluency assessment, selected one output for post-editing, and justified their deci-sion in written reports. Descriptive counts are reported for all 23 projects, while qualitative interpretation is based on the 22 cases accompanied by written reports. Results show that students did not treat automatic metrics as final authority: final post-editing selections often diverged from metric rankings and were justified through adequacy, fluency, terminology, naturalness, and expected post-editing ef-fort. The study therefore does not bench-mark systems under controlled conditions; it analyses how students justified system choice within an authentic classroom as-signment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。