AI生成的试题更难,学生却分不清真假。
LLMs in the Classroom: Outcomes and Perceptions of Questions Written with the Aid of AI
- 用SBERT比对AI与人工出题的文本相似度
- 学生答AI题得分低近9%,差异显著
- 适合教育研究者和课程设计者参考
我们随机部署了由AI工具生成和人工编写的题目,评估学生正确作答的能力以及辨别题目来源的能力。通过SBERT计算人类与ChatGPT生成题目的代表性向量,并与课程教材进行余弦相似度比较。非显著的Mann-Whitney U检验(z = 1.018, p = .309)表明,学生无法感知题目是否由AI辅助生成。然而,学生在AI生成题目上的得分比人工题低约9%(z = 2.702, p < .01),可能因AI题更难或学生更熟悉教师出题风格所致。研究提示:尽管可用LLM辅助命题,但仍需确保题目公平、合理且贴合课程内容。
原文摘要 · Abstract (English)
We randomly deploy questions constructed with and without use of the LLM tool and gauge the ability of the students to correctly answer, as well as their ability to correctly perceive the difference between human-authored and LLM-authored questions. In determining whether the questions written with the aid of ChatGPT were consistent with the instructor's questions and source text, we computed representative vectors of both the human and ChatGPT questions using SBERT and compared cosine similarity to the course textbook. A non-significant Mann-Whitney U test (z = 1.018, p = .309) suggests that students were unable to perceive whether questions were written with or without the aid of ChatGPT. However, student scores on LLM-authored questions were almost 9% lower (z = 2.702, p < .01). This result may indicate that either the AI questions were more difficult or that the students were more familiar with the instructor's style of questions. Overall, the study suggests that while there is potential for using LLM tools to aid in the construction of assessments, care must be taken to ensure that the questions are fair, well-composed, and relevant to the course material.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。