用大模型批量检查课程材料,识别AI滥用风险,助力高校质量保障。
Scalable Classification of Course Information Sheets Using Large Language Models: A Reusable Institutional Method for Academic Quality Assurance
- 构建四阶段流程,通过迭代优化提示词实现精准分类
- 扫描4684份课程资料,60.3%被标记为高风险,模型准确率达87%
- 方法可复用至可持续性、无障碍等审计场景,适合教育管理者
高等教育机构面临加强课程设计中生成式AI(GenAI)整合审计的压力。本文提出一个端到端的大语言模型(LLM)方法,可大规模扫描课程信息表,识别学生可能滥用GenAI的评估环节,通过迭代优化验证系统性能,并通过直接沟通将结果落地。研究构建了四阶段流程:(0) 手动试点采样;(1) 多模型对比下的迭代提示工程;(2) 对弗莱堡大学(VUB)2024-2025学年4,684份本科与硕士课程信息表进行全量扫描,采用三级风险分类体系(明确风险、潜在风险、低风险),自动生成报告并邮件发送至教学团队(91.4%地址匹配成功);(3) 在下一版目录发布后对4,675份文件进行纵向重扫。经五轮提示词优化,模型与专家标注一致率达87%。GPT-4o因在实习与实践类模糊案例中表现更优被选为生产模型。首年扫描结果显示60.3%课程为明确风险,15.2%为潜在风险,24.5%为低风险。第二年对比显示风险分布显著变化,实践导向项目改进最为明显。该方法使机构能快速将异构目录数据转化为结构化、可操作的智能信息,具备向可持续性、无障碍、教学一致性等审计领域迁移的能力,为高校治理中负责任的LLM应用提供模板。
原文摘要 · Abstract (English)
Purpose: Higher education institutions face increasing pressure to audit course designs for generative AI (GenAI) integration. This paper presents an end-to-end method for using large language models (LLMs) to scan course information sheets at scale, identify where assessments may be vulnerable to student use of GenAI tools, validate system performance through iterative refinement, and operationalise results through direct stakeholder communication and effort. Method: We developed a four-phase pipeline: (0) manual pilot sampling, (1) iterative prompt engineering with multi-model comparison, (2) full production scan of 4,684 Bachelor and Master course information sheets (Academic Year 2024-2025) from the Vrije Universiteit Brussel (VUB) with automated report generation and email distribution to teaching teams (91.4% address-matched) using a three-tier risk taxonomy (Clear risk, Potential risk, Low risk), and (3) longitudinal re-scan of 4,675 sheets after the next catalogue release. Results: Five iterations of prompt refinement achieved 87% agreement with expert labels. GPT-4o was selected for production based on superior handling of ambiguous cases involving internships and practical components. The Year 1 scan classified 60.3% of courses as Clear risk, 15.2% as Potential risk, and 24.5% as Low risk. Year 2 comparison revealed substantial shifts in risk distributions, with improvements most pronounced in practice-oriented programmes. Implications: The method enables institutions to rapidly transform heterogeneous catalogue data into structured and actionable intelligence. The approach is transferable to other audit domains (sustainability, accessibility, pedagogical alignment) and provides a template for responsible LLM deployment in higher education governance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。