用布鲁姆分类法评估大模型教书能力,发现它能提升难度却难降级。
From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs

- 以布鲁姆分类为尺度,量化模型调整题目认知难度的能力。
- 两模型都能提高任务难度,但降低难度时表现差,存在方向性失衡。
- 适合教育技术研究者和想提升模型教学能力的开发者参考。
我们提出一个与布鲁姆分类法对齐的框架,用于衡量大语言模型(LLMs)在教育中的控制能力:即在保持任务教学意图的同时,将其认知需求调整至指定学习目标。以计算机科学教育中的编程任务为例,研究解题与适配学习者之间的差距。采用修订版布鲁姆分类法作为认知需求的操作化量表,评估两种干预设置:通用难度控制(使任务更难或更易)和布鲁姆层级控制(针对更高或更低的布鲁姆层级)。对比两个匹配的Qwen3-Next模型对:Qwen3-Next-80B-A3B-Instruct与Qwen3-Coder-Next,共在三个基准上的2,520个任务上进行评估。结果揭示出显著的方向不对称性:两个模型均能可靠提升认知需求,但在降低需求方面表现不佳。通过语义差聚类和层间费希尔判别比探测进一步刻画结果。在可控对比中,通用模型在中间层对两类对比均表现出更强可分性;而编码专用模型在通用难度对比中可分性较弱,在布鲁姆层级对比中出现更深峰值。表明强执行能力并不自动带来符合布鲁姆目标的教育控制能力。
原文摘要 · Abstract (English)
We introduce a Bloom-aligned framework for measuring educational control in Large Language Models (LLMs): the ability to preserve a task's instructional intent while shifting its cognitive demand toward specified learning objectives. We apply this framework to programming tasks in computer science education to study the gap between solving tasks and adapting them for learners. Using revised Bloom's Taxonomy as an operational scale of cognitive demand, we evaluate two intervention settings: general difficulty control, where models are asked to make tasks harder or easier, and Bloom's control, where models are asked to target higher or lower Bloom's levels. We evaluate a matched Qwen3-Next model pair, comparing Qwen3-Next-80B-A3B-Instruct with Qwen3-Coder-Next across 2,520 tasks from three benchmarks. The framework reveals a robust directional asymmetry: both models reliably increase cognitive demand, but struggle to lower it. We further characterize these outcomes with semantic-delta clustering and layer-wise Fisher's Discriminant Ratio probing. Within this controlled comparison, the general model shows clearer middle-layer separability for both general difficulty and Bloom-control contrasts, whereas the coder model shows weaker separability for general difficulty and a deeper peak for Bloom-control contrasts. These results show that strong execution performance does not automatically entail Bloom-aligned educational control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。