arXiv:2602.08796cs.AIcs.CL2026-02

用AI生成阅读理解题的属性矩阵,效果比人类专家还高。

The Use of AI Tools to Develop and Validate Q-Matrices

  • 让AI基于教材文本生成题目标签矩阵
  • 谷歌Gemini 2.5 Pro与标准矩阵吻合度最高(Kappa=0.63)
  • 新版本AI反而表现下降,提示需持续评估

构建Q矩阵是认知诊断模型中的关键但耗时环节。本研究探讨通用语言模型能否辅助Q矩阵开发,将AI生成的矩阵与Li和Suen(2013)验证过的阅读理解测试Q矩阵进行对比。2025年5月,多个AI模型使用与人类专家相同的训练材料,通过Cohen's kappa评估AI生成矩阵、验证矩阵及人工标注矩阵的一致性。结果显示各AI模型间差异显著,其中Google Gemini 2.5 Pro与标准矩阵一致性最高(Kappa = 0.63),优于所有人类专家。然而2026年1月的后续分析显示,新版本AI模型与标准矩阵的一致性下降。研究讨论了其启示与未来方向。

原文摘要 · Abstract (English)

Constructing a Q-matrix is a critical but labor-intensive step in cognitive diagnostic modeling (CDM). This study investigates whether AI tools (i.e., general language models) can support Q-matrix development by comparing AI-generated Q-matrices with a validated Q-matrix from Li and Suen (2013) for a reading comprehension test. In May 2025, multiple AI models were provided with the same training materials as human experts. Agreement among AI-generated Q-matrices, the validated Q-matrix, and human raters' Q-matrices was assessed using Cohen's kappa. Results showed substantial variation across AI models, with Google Gemini 2.5 Pro achieving the highest agreement (Kappa = 0.63) with the validated Q-matrix, exceeding that of all human experts. A follow-up analysis in January 2026 using newer AI versions, however, revealed lower agreement with the validated Q-matrix. Implications and directions for future research are discussed.

认知诊断AI辅助Q矩阵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。