用大模型评估阅读理解题的认知难度,效果不错但自知力不足。
Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?
- 聚焦证据范围与转换层级,量化题目推理复杂度
- 大模型预测认知难度相关性达0.72,接近人工标注水平
- 适合教育评测、题库建设者参考,但需警惕模型误判
评估阅读理解题的认知复杂度对预判题目难度至关重要。传统上依赖人工标注的推理过程特征,如证据范围和转换层级,难以通过现有NLP工具提取。本研究探讨大语言模型(LLMs)是否可基于这两个维度估算题目认知负担。实验表明,LLMs能有效逼近人工标注的认知复杂度,相关性达0.72。进一步分析发现,尽管模型能正确作答,却常无法识别自身推理所依赖的关键特征,反映出其推理能力与元认知意识之间的差距。
原文摘要 · Abstract (English)
Estimating the cognitive complexity of reading comprehension (RC) items is crucial for assessing item difficulty before it is administered to learners. Unlike syntactic and semantic features, such as passage length or semantic similarity between options, cognitive features that arise during answer reasoning are not readily extractable using existing NLP tools and have traditionally relied on human annotation. In this study, we examine whether large language models (LLMs) can estimate the cognitive complexity of RC items by focusing on two dimensions-Evidence Scope and Transformation Level-that indicate the degree of cognitive burden involved in reasoning about the answer. Our experimental results demonstrate that LLMs can approximate the cognitive complexity of items, indicating their potential as tools for prior difficulty analysis. Further analysis reveals a gap between LLMs' reasoning ability and their metacognitive awareness: even when they produce correct answers, they sometimes fail to correctly identify the features underlying their own reasoning process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。