研究生难区分真人与AI写作,提示需改进训练方法
Can postgraduate translation students identify machine-generated text?
- 23名翻译研究生经短期培训后判断文本来源
- 平均准确率低,仅2人表现突出,普遍误判
- 发现低突发性与自相矛盾更倾向AI生成
随着生成式人工智能广泛用于多语言内容创作并绕过传统翻译方式,本研究考察语言训练背景的个体识别机器生成文本(ST)与人类写作(HT)的能力。23名研究生在简短培训后分析意大利散文片段,并对文本来源的可能性进行评分。结果显示,平均而言学生难以区分两者,仅两人表现显著准确。深入分析表明,学生常在两类文本中识别出相同异常特征,但低突发性与自相矛盾更常关联于机器生成文本。这说明当前训练方法需优化,同时引发对是否仍需编辑使AI文本更自然的思考。
原文摘要 · Abstract (English)
Given the growing use of generative artificial intelligence as a tool for creating multilingual content and bypassing both machine and traditional translation methods, this study explores the ability of linguistically trained individuals to discern machine-generated output from human-written text (HT). After brief training sessions on the textual anomalies typically found in synthetic text (ST), twenty-three postgraduate translation students analysed excerpts of Italian prose and assigned likelihood scores to indicate whether they believed they were human-written or AI-generated (ChatGPT-4o). The results show that, on average, the students struggled to distinguish between HT and ST, with only two participants achieving notable accuracy. Closer analysis revealed that the students often identified the same textual anomalies in both HT and ST, although features such as low burstiness and self-contradiction were more frequently associated with ST. These findings suggest the need for improvements in the preparatory training. Moreover, the study raises questions about the necessity of editing synthetic text to make it sound more human-like and recommends further research to determine whether AI-generated text is already sufficiently natural-sounding not to require further refinement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。