测试AI对成语的评分能力,发现谷歌Gemini表现最佳。
Can generative AI figure out figurative language? The influence of idioms on essay scoring by ChatGPT, Gemini, and Deepseek
- 用含成语和不含成语的作文对比测试三款AI评分
- Gemini与人工评分一致性最高,处理比喻语言更像人
- 无群体偏差,适合未来独立用于作文评分
生成式AI在自动评阅学生作文方面展现出潜力。本研究基于语料库语言学与计算语言学视角,评估了三款生成式AI模型(ChatGPT、Gemini、Deepseek)在含成语与不含成语作文上的评分表现。从348篇学生作文中构建两组等长作文集:一组每篇含多个成语,另一组无成语。三模型各对两组作文重复评分三次,使用与人工评分一致的量表。结果显示所有模型评分一致性良好,其中Gemini在与人工评分者的一致性上最优;含成语作文中,Gemini的评分模式最接近人类。且各模型均未表现出对任何人口群体的可检测偏差。尽管存在混合使用潜力,但鉴于其对修辞语言的处理能力,Gemini是最适合作为独立作文评分工具的候选者。
原文摘要 · Abstract (English)
The developments in Generative AI technologies have paved the way for numerous innovations in different fields. Recently, Generative AI has been proposed as a competitor to AES systems in evaluating student essays automatically. Considering the potential limitations of AI in processing idioms, this study assessed the scoring performances of Generative AI models for essays with and without idioms by incorporating insights from Corpus Linguistics and Computational Linguistics. Two equal essay lists were created from 348 student essays taken from a corpus: one with multiple idioms present in each essay and another with no idioms in essays. Three Generative AI models (ChatGPT, Gemini, and Deepseek) were asked to score all essays in both lists three times, using the same rubric used by human raters in assigning essay scores. The results revealed excellent consistency for all models, but Gemini outperformed its competitors in interrater reliability with human raters. There was also no detectable bias for any demographic group in AI assessment. For essays with multiple idioms, Gemini followed a the most similar pattern to human raters. While the models in the study demonstrated potential for a hybrid approach, Gemini was the best candidate for the task due to its ability to handle figurative language and showed promise for handling essay-scoring tasks alone in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。