测试ChatGPT解韩语数学题表现,准确率近67%。
On the robustness of ChatGPT in teaching Korean Mathematics
- 用586道韩语数学题评估ChatGPT性能,验证其在非英语环境下的能力。
- 正确解答391题,准确率达66.72%,在分类任务中表现良好。
- 适合教育研究者、多语言AI开发者关注其语言偏见与优化方向。
ChatGPT作为人工智能模型,有潜力改变教育方式。但其解决非英语问题的效果仍不确定。本研究使用586道韩语数学题评估ChatGPT的鲁棒性,结果显示其准确率为66.72%,共正确回答391题。我们还评估了其对数学题目的评分能力,涵盖十一项标准,并进行主题分析。结果表明,ChatGPT的评分与教育理论及考生视角高度一致。尽管在题目分类上表现良好,但在非英语语境下仍存在挑战,凸显改进空间。未来研究应关注语言偏见问题,通过领域特定优化和多语言训练提升跨语言准确性,助力个性化教育发展。
原文摘要 · Abstract (English)
ChatGPT, an Artificial Intelligence model, has the potential to revolutionize education. However, its effectiveness in solving non-English questions remains uncertain. This study evaluates ChatGPT's robustness using 586 Korean mathematics questions. ChatGPT achieves 66.72% accuracy, correctly answering 391 out of 586 questions. We also assess its ability to rate mathematics questions based on eleven criteria and perform a topic analysis. Our findings show that ChatGPT's ratings align with educational theory and test-taker perspectives. While ChatGPT performs well in question classification, it struggles with non-English contexts, highlighting areas for improvement. Future research should address linguistic biases and enhance accuracy across diverse languages. Domain-specific optimizations and multilingual training could improve ChatGPT's role in personalized education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。