测试德国AI作业评分工具,发现其评分随意且无法识别作弊。
Chatbots im Schulunterricht: Wir testen das Fobizz-Tool zur automatischen Bewertung von Hausaufgaben
- 用两轮测试评估Fobizz AI评分工具的实用性
- 仅用ChatGPT生成文本才能获最高分,错误内容常被忽略
- 适合关注AI教育应用风险的教师与政策制定者
本研究考察德国公司Fobizz推出的AI评分工具「AI Grading Assistant」,旨在辅助教师批改学生作业并提供反馈。在教育系统负担过重、社会对AI解决教育问题寄予厚望的背景下,通过两轮测试评估该工具的功能适用性。结果表明:工具给出的分数和评语往往随机,即使采纳其建议也未改善;仅当提交由ChatGPT生成的文本时才能获得高分;虚假陈述和无意义内容常被漏检;部分评分标准的执行不一致且缺乏透明度。这些缺陷源于大语言模型(LLMs)的固有局限,短期内难以根本改进。研究批判了将AI视为系统性教育问题速解方案的趋势,认为Fobizz将其宣传为客观、省时的解决方案是误导且不负责任的。最后呼吁对教育场景中使用AI工具进行系统性评估和学科特定的教学法审查。
原文摘要 · Abstract (English)
This study examines the AI-powered grading tool "AI Grading Assistant" by the German company Fobizz, designed to support teachers in evaluating and providing feedback on student assignments. Against the societal backdrop of an overburdened education system and rising expectations for artificial intelligence as a solution to these challenges, the investigation evaluates the tool's functional suitability through two test series. The results reveal significant shortcomings: The tool's numerical grades and qualitative feedback are often random and do not improve even when its suggestions are incorporated. The highest ratings are achievable only with texts generated by ChatGPT. False claims and nonsensical submissions frequently go undetected, while the implementation of some grading criteria is unreliable and opaque. Since these deficiencies stem from the inherent limitations of large language models (LLMs), fundamental improvements to this or similar tools are not immediately foreseeable. The study critiques the broader trend of adopting AI as a quick fix for systemic problems in education, concluding that Fobizz's marketing of the tool as an objective and time-saving solution is misleading and irresponsible. Finally, the study calls for systematic evaluation and subject-specific pedagogical scrutiny of the use of AI tools in educational contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。