对比人类与AI在多模态理科题上的表现,发现AI在视觉题上明显落后。
Challenges for AI in Multimodal STEM Assessments: a Human-AI Comparison
- 构建201道大学级多模态理科学术题数据集,标注图像类型、复杂度等特征。
- 最优模型仅答对58.5%题目,人类在含图题中持续领先。
- 提醒教育者设计难倒AI但不增加学生负担的题目以维护学术诚信。
生成式AI系统快速演进,多模态输入能力使其能处理超越文本的任务。在教育领域,这些进展可能重塑评估设计与答题方式,带来机遇与挑战。为此,我们构建了一个高质量数据集,包含201道大学级别STEM问题,人工标注了图像类型、作用、问题复杂度及题型等特征。研究分析这些特征如何影响生成式AI性能与学生表现。我们评估了四种模型家族与五种提示策略,并将结果与每道题平均546名学生的回答进行对比。尽管最佳模型通过多数投票平均答对58.5%的问题,人类在涉及视觉成分的问题上始终优于AI。值得注意的是,人类表现受学科影响但对题型特征稳定,而AI表现同时受学科和题型特征显著影响。最后,我们为教育者提供可操作建议:通过设计挑战当前AI系统的题目特征,在不增加学生认知负担的前提下提升学术完整性。
原文摘要 · Abstract (English)
Generative AI systems have rapidly advanced, with multimodal input capabilities enabling reasoning beyond text-based tasks. In education, these advancements could influence assessment design and question answering, presenting both opportunities and challenges. To investigate these effects, we introduce a high-quality dataset of 201 university-level STEM questions, manually annotated with features such as image type, role, problem complexity, and question format. Our study analyzes how these features affect generative AI performance compared to students. We evaluate four model families with five prompting strategies, comparing results to the average of 546 student responses per question. Although the best model correctly answers on average 58.5 % of the questions using majority vote aggregation, human participants consistently outperform AI on questions involving visual components. Interestingly, human performance remains stable across question features but varies by subject, whereas AI performance is susceptible to both subject matter and question features. Finally, we provide actionable insights for educators, demonstrating how question design can enhance academic integrity by leveraging features that challenge current AI systems without increasing the cognitive burden for students.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。