arXiv:2510.06411cs.CL2025-10被引 1

用大模型生成符合教学目标的实验题,让虚拟实验更贴合课堂需求。

Instructional Goal-Aligned Question Generation for Student Evaluation in Virtual Lab Settings: How Closely Do LLMs Actually Align?

  • 通过师生对话理解教学目标,结合实验知识结构生成问题。
  • 优化提示词后,问题格式正确率超90%,质量提升0.8分(李克特量表)。
  • 适合教育工作者和教育AI研发者,尤其关注个性化习题生成。

虚拟实验为探究式科学学习提供了宝贵机会,但教师常难以将其适配至教学目标。第三方资源可能与课堂需求不符,自建资源又耗时且难规模化。大语言模型(LLMs)为此提供了新路径。本文提出一种教学目标对齐的问题生成框架,支持教师通过自然语言交互,利用LLMs生成与实验仿真一致、具有教学意义的问题。框架包含四个模块:基于师生对话的教学目标理解、基于知识单元与关系分析的实验理解、用于构建认知与教学意图的问题分类体系,以及控制提示细节的TELeR分类体系。早期设计参考小规模教师辅助案例研究,最终评估涵盖19个开源LLM生成的1100+道问题。在教学目标与实验理解双重引导下,问题分类体系使认知要求提升(开放题与关系类问题质量提高0.29–0.39分),优化后的TELeR提示显著提升格式可解析性(80%)、格式遵循度(>90%)。大模型表现最优:可解析性提升37.1%,遵循度提升25.7%,平均质量提升0.8李克特量表点。

原文摘要 · Abstract (English)

Virtual Labs offer valuable opportunities for hands-on, inquiry-based science learning, yet teachers often struggle to adapt them to fit their instructional goals. Third-party materials may not align with classroom needs, and developing custom resources can be time-consuming and difficult to scale. Recent advances in Large Language Models (LLMs) offer a promising avenue for addressing these limitations. In this paper, we introduce a novel alignment framework for instructional goal-aligned question generation, enabling teachers to leverage LLMs to produce simulation-aligned, pedagogically meaningful questions through natural language interaction. The framework integrates four components: instructional goal understanding via teacher-LLM dialogue, lab understanding via knowledge unit and relationship analysis, a question taxonomy for structuring cognitive and pedagogical intent, and the TELeR taxonomy for controlling prompt detail. Early design choices were informed by a small teacher-assisted case study, while our final evaluation analyzed over 1,100 questions from 19 open-source LLMs. With goal and lab understanding grounding questions in teacher intent and simulation context, the question taxonomy elevates cognitive demand (open-ended formats and relational types raise quality by 0.29-0.39 points), and optimized TELeR prompts enhance format adherence (80% parsability, >90% adherence). Larger models yield the strongest gains: parsability +37.1%, adherence +25.7%, and average quality +0.8 Likert points.

虚拟实验大模型教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。