用大模型推理过程中的认知片段预测题目难度,更准且可解释。
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction

- 将大模型推理轨迹拆解为认知事件序列,捕捉解题过程动态
- 在4个真实数据集上超越基线,SAT任务相对提升8.1%
- 适合教育评估、智能命题与可解释性研究者
预测人类题目难度是教育测评的核心,可靠估计有助于公平性和高效测验设计。现有方法多依赖昂贵的人工校准或题目文本表征,难以揭示使题目变难的认知机制。我们认为难度不仅是题目文本属性,更是问题求解负担的可观测结果。大推理模型(LRM)通过推理轨迹提供可扩展的过程证据,但需结构化以支持可解释建模。为此,我们提出Epi2Diff框架,将LRM推理轨迹映射为具认知基础的事件序列。这些事件将轨迹分段为功能性解题状态,使难度可由推理尺度、努力分配和状态转换建模。Epi2Diff提取紧凑的事件动态特征,并结合语义题目表征进行人类难度预测。在四个真实世界人类难度数据集上的实验表明,Epi2Diff持续优于强基线,包括微调小模型、LLM上下文学习和监督式LLM适配。在基于SAT的分类基准上,相比监督式LLM微调基线平均相对提升8.1%。进一步分析显示,难题引发更费力、迭代性更强、以实现为中心的事件动态,而非仅更长的回答。结果表明,大模型推理轨迹中的认知事件为人类题目难度提供了具有预测力和可解释性的过程表征,为推理模型驱动的教育测量提供了新视角。
原文摘要 · Abstract (English)
Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction. Existing methods often depend on costly human calibration or item-level textual representations, providing limited evidence about the cognitive processes that make items difficult. We argue that difficulty should be viewed not only as a property of item text, but also as an observable consequence of the problem-solving burden an item induces. Large Reasoning Models (LRMs) offer scalable process evidence through reasoning traces, but such evidence must be structured to support interpretable modeling. To this end, we introduce Epi2Diff (Episode to Difficulty), a framework that maps LRM reasoning traces into cognitively grounded episode sequences. These episodes group trace segments into functional problem-solving states, enabling difficulty to be modeled through reasoning scale, effort allocation, and state transitions. Epi2Diff extracts compact episode-dynamic features and combines them with semantic item representations for human difficulty prediction. Experiments on four real-world human difficulty datasets show that Epi2Diff consistently outperforms strong baselines, including fine-tuned small language models, LLM in-context learning, and supervised LLM adaptation. On SAT-derived classification benchmarks, Epi2Diff achieves an 8.1% average relative gain over supervised LLM fine-tuning baselines. Further analyses show that harder items induce more effortful, iterative, and implementation-centered episode dynamics, rather than merely longer responses. These results demonstrate that cognitive episodes in LRM reasoning traces provide a predictive and interpretable process representation for human item difficulty, offering a new lens for educational measurement with reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。