用多智能体大模型模拟自适应测试,真实还原答题与选题互动。
AgentCAT: Simulating Computerized Adaptive Testing via Multi-Agent Large Language Models

- 构建三智能体系统:考生、选题、监督者协同模拟动态评估过程。
- 在真实数据集上实现能力估计收敛,选题兼顾难度与知识点覆盖。
- 适合教育科技研究者、个性化学习系统开发者参考。
计算机化自适应测试(CAT)是个性化教育的关键技术,旨在通过动态匹配当前能力估计的题目来精准评估考生水平。然而,现有研究受限于静态离线数据和单一模块优化,常将动态评估简化为静态序列预测,且聚焦于选题或诊断等孤立环节,忽视整体交互过程。为此,我们提出AgentCAT,一个基于大语言模型的多智能体仿真系统,用于构建高保真动态测试基准环境。该框架包含三个模块:(1) 考生智能体结合记忆检索与思维链推理,根据认知画像模拟作答;(2) 选题智能体采用粗粒度分桶与知识图谱探索,平衡局部难度与全局覆盖;(3) 监督智能体通过双重审核与鲁棒更新机制保障收敛性与有效性。我们在两个真实数据集上从宏观能力收敛、微观交互逻辑和数据稀疏鲁棒性三个维度验证框架效果。结果表明,AgentCAT实现了有效的能力估计,其选题策略在难度适应与教学连贯性间取得良好平衡,符合人类教学直觉。
原文摘要 · Abstract (English)
Computerized Adaptive Testing (CAT), as a key technology for personalized education, aims to accurately assess examinee proficiency by retrieving exercises dynamically matching current ability estimates. However, existing CAT research is constrained by limitations of static offline data and isolated component optimization. Restricted by partial labels in offline logs, researchers degrade the dynamic assessment process into static sequence prediction. Current research focuses on isolated perspectives, e.g., selection or diagnosis, neglecting the overall CAT interaction process. To address this, we propose AgentCAT, a Large Language Model-based multi-agent simulation system, to construct a high-fidelity benchmarking environment for dynamic testing. This framework comprises three modules: (1) The examinee agent with memory retrieval and Chain-of-Thought reasoning simulates responses based on cognitive profiles; (2) The selection agent uses coarse-to-fine bucketing and knowledge graph exploration to balance local difficulty and global coverage; (3) The supervisor uses dual-auditing and robust update to ensure convergence and validity. To validate the framework, we evaluated on two real-world datasets across three dimensions: macro-level ability convergence, micro-level interaction logic, and data sparsity resilience. Results show AgentCAT achieves effective ability estimation, and its selection strategy balances difficulty adaptation and instructional coherence, aligning with human pedagogical intuition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。