arXiv:2604.16398cs.CYcs.AI2026-04中稿 · AIED 2026

用AI生成+模型评估,让专家快速优化知识矩阵。

A Framework for Human-AI Q-Matrix Refinement: A NeuralCDM Evaluation

论文配图:A Framework for Human-AI Q-Matrix Refinement: A NeuralCDM Evaluation
图 1 · 摘自论文原文
  • LLM通过有误解意识的提示生成候选知识矩阵
  • 迭代优化后模型拟合度提升(AUC 0.780 > 0.717)
  • 本地部署模型效果媲美云端,适合隐私敏感场景

Q矩阵是理论驱动评估与学习分析的核心,能明确揭示题目需求、学生知识成分及误解。但传统由专家手工构建的Q矩阵耗时长、易主观且难验证。本文提出一种人机协同的Q矩阵优化框架:大语言模型(LLMs)通过结构化、包含误解意识的提示生成候选矩阵,神经认知诊断模型(NeuralCDM)则基于学生作答数据评估候选矩阵的解释能力。我们在热力学测评数据集上应用该框架,并对比本地部署与云端调用的LLMs表现。结果表明,经迭代优化的LLM生成矩阵在模型拟合度上超越专家基准(AUC 0.780 vs. 0.717),且本地部署模型性能可比云端服务,支持隐私保护型部署。

原文摘要 · Abstract (English)

Q-matrices are a cornerstone of theory-driven assessment and learning analytics, making item demands and students' underlying knowledge components and misconceptions explicit and actionable. However, Q-matrices are typically crafted by experts, making them time-consuming to build, prone to subjectivity, and difficult to validate empirically. We propose a framework for human-AI Q-matrix refinement in which large language models (LLMs) generate candidate Q-matrices using structured, misconception-aware prompting, and NeuralCDM provides an empirical evaluation layer to compare candidates based on how well they explain student response data. We apply the framework to a thermodynamics assessment dataset and benchmark locally deployed LLMs against cloud-served models. Results show that iteratively refined LLM-generated Q-matrices can exceed expert-baseline model fit (AUC 0.780 vs. 0.717), and that locally deployed models achieve comparable performance to cloud APIs, supporting privacy-preserving deployment.

认知诊断人机协作LLM应用隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。