arXiv:2606.13684cs.CYcs.AI2026-06中稿 · AIED 2026

用提示工程让大模型跨数据集准确分类教育题目认知层次。

Cross-Dataset Bloom Question Classification: Supervised Models and Prompted LLMs

论文配图:Cross-Dataset Bloom Question Classification: Supervised Models and Prompted LLMs
图 1 · 摘自论文原文
  • 用上下文示例+课程动词组合提示,提升大模型分类效果。
  • 传统模型跨数据集性能下降明显,大模型更稳定可靠。
  • 适合需要批量处理试题的认知层次分类的教师和教育工具开发者。

自动对评估题进行布卢姆分类可显著减轻教师负担,但标注主观性强且依赖教师。以往机器学习与深度学习方法在单一数据集上表现优异,却极少在跨数据集场景下评估,真实泛化能力不明;同时大模型在布卢姆分类中的有效性尚未系统研究。本文评估了现有机器学习/深度学习模型在跨数据集上的表现,并测试了多种提示策略的大模型在五个数据集上的表现;最优提示策略结合了上下文示例与课程特定动作动词。监督式模型在未见数据集上性能大幅下降,而大模型表现更稳定,表明其在多元教育场景中具有更强鲁棒性。基于最优提示策略,我们还设计了一款轻量级用户界面,支持教师自动分类大量试题库;可用性研究表明该工具工作负荷低、易用性高。

原文摘要 · Abstract (English)

Automatic Bloom's taxonomy classification of assessment questions can substantially reduce instructor workload, but labeling is subjective and teacher-dependent. Prior machine learning (ML) and deep learning (DL) approaches reported strong within-dataset results, yet were rarely evaluated in cross-dataset settings, leaving real-world generalizability unclear; meanwhile, LLM effectiveness for Bloom question classification has not been systematically studied. We evaluated the cross-dataset generalization of existing ML/DL methods and assessed LLMs with multiple prompting strategies on five datasets; the best prompting strategy combined in-context examples with course-specific action verbs. Supervised ML/DL models degraded substantially on unseen datasets, whereas LLMs were more stable, suggesting a robust alternative across diverse educational contexts. Based on the best prompting strategy, we also presented a lightweight UI that supports instructors in automatically classifying large question banks; a usability study indicated low workload and high usability.

教育AI提示工程大模型布卢姆分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。