用强化学习建模答题过程,更高效且可解释。
Reinforcement Learning Measurement Model

- 用共享参数函数分离个体敏感度与任务价值,提升计算效率
- 在跳棋模拟中准确率更高、耗时更低,复杂度越高优势越明显
- 适合研究真实行为决策过程的教育评估与心理测量
交互式测评生成序列过程数据,传统项目反应模型难以处理。现有基于马尔可夫决策过程的测量方法(如MDP-MM)虽将行为选择与状态-动作值关联,但依赖个体特定的表格型价值函数,难以扩展至大规模任务。本文提出强化学习测量模型(RLMM),通过共享参数化动作价值函数,解耦个体选择敏感度与任务价值表示,显著提升大规模过程数据下的估计效率。模型结合Boltzmann选择规则与归一化优势、软贝尔曼一致性惩罚,以及块坐标最大后验估计法实现联合推断,并能提供步骤级影响诊断,识别关键决策点。在跳棋模拟中,RLMM的估计精度更高,运行时间显著降低,且随任务复杂度提升优势扩大;在AQUALAB游戏日志中,个体参数与累积奖励、任务完成度及行为效率呈正相关。结果表明,该模型可将基于决策过程的心理测量扩展至更大、更真实的环境,同时保持与决策步骤可解释的潜在特质关联。
原文摘要 · Abstract (English)
Interactive assessments generate sequential process data that are not well handled by conventional item response models. Existing MDP-based measurement approaches, such as the Markov decision process measurement model (MDP-MM, LaMar, 2018), link action choices to state-action values, but their reliance on person-specific tabular value functions makes them difficult to scale beyond small, fully enumerated tasks. We propose the Reinforcement Learning Measurement Model (RLMM), a measurement framework that decouples person-level choice sensitivity from task-level value representation through a shared parametric action-value function, making estimation more computationally efficient for larger process-data settings. The model combines a Boltzmann choice rule with normalized advantages, a soft Bellman consistency penalty, and a block-coordinate MAP procedure for joint estimation, while also yielding step-level influence diagnostics for identifying behaviorally critical decisions. In peg-solitaire simulations, the RLMM achieved higher estimation accuracy and substantially lower runtime than the original MDP-MM, with advantages increasing as task complexity grew. In AQUALAB gameplay logs, the estimated person parameter was positively associated with cumulative reward, task completion, and behavioral efficiency. These results show that the RLMM extends decision-process-based psychometric models to larger and more behaviorally realistic environments while preserving an interpretable latent trait tied to decision making steps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。