用大模型自动评估课程,能精准分析教学细节并给出改进建议。
An Exploration of Higher Education Course Evaluation by Large Language Models
- 用大模型分析课堂对话和课程资料,实现细粒度评估
- 微调后的Llama模型评分更区分度高,与专家评价相关性强
- 适合高校教务部门做大规模教学质量监控
课程评价在高等教育中对保障教学质量、指导课程建设至关重要。传统方法如学生问卷、课堂观察和专家评审存在主观性强、人力成本高、难以扩展等问题。随着大语言模型(LLMs)的发展,自动化、精细化、可扩展的课程评价成为可能。本研究利用来自中国某高校的100门课程数据,结合课堂互动记录,评估三种代表性LLM在微观(课堂讨论分析)和宏观(整体课程评估)层面的表现。结果表明,LLMs能有效提取关键教学特征,并生成与专家判断一致的结构化评价。其中,经过微调的Llama模型表现最优,其评分分布更具区分度,且与人工评价相关性更强。主要发现包括:(1) LLM可在微观与宏观层面可靠完成系统化、可解释的课程评估;(2) 微调与提示工程显著提升评估准确性与一致性;(3) 模型生成的反馈为教学改进提供可操作建议。这些成果展示了基于大模型的评价在大规模高等教育质量保障与决策中的应用潜力。
原文摘要 · Abstract (English)
Course evaluation plays a critical role in ensuring instructional quality and guiding curriculum development in higher education. However, traditional evaluation methods, such as student surveys, classroom observations, and expert reviews, are often constrained by subjectivity, high labor costs, and limited scalability. With recent advancements in large language models (LLMs), new opportunities have emerged for generating consistent, fine-grained, and scalable course evaluations. This study investigates the use of three representative LLMs for automated course evaluation at both the micro level (classroom discussion analysis) and the macro level (holistic course review). Using classroom interaction transcripts and a dataset of 100 courses from a major institution in China, we demonstrate that LLMs can extract key pedagogical features and generate structured evaluation results aligned with expert judgement. A fine-tuned version of Llama shows superior reliability, producing score distributions with greater differentiation and stronger correlation with human evaluators than its counterparts. The results highlight three major findings: (1) LLMs can reliably perform systematic and interpretable course evaluations at both the micro and macro levels; (2) fine-tuning and prompt engineering significantly enhance evaluation accuracy and consistency; and (3) LLM-generated feedback provides actionable insights for teaching improvement. These findings illustrate the promise of LLM-based evaluation as a practical tool for supporting quality assurance and educational decision-making in large-scale higher education settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。