用多智能体LLM从游戏行为中无声评估财务素养,效果显著优于传统方法。
Agentic Knowledge Tracing: A Multi-Agent LLM Architecture for Stealth Assessment of Financial Literacy in Serious Games

- 分领域智能体协作分析玩家行为轨迹,实现精细化能力追踪。
- 与学习进步和测试成绩显著相关(r=0.276~0.333),预测试无关联。
- 适合教育游戏开发者、认知评估研究者及财务素养教学设计者。
在不干扰学习体验的前提下,对严肃游戏中的财务素养进行评估仍是教育领域的关键挑战。本文提出Agentic BKT流水线,一种基于多智能体大语言模型架构的隐蔽式财务能力评估方法。该系统通过四阶段处理:(1) 游戏将每个玩家决策记录为结构化事件日志;(2) 使用LLM事件分类器对每项操作进行四点量表标注,经三位领域专家验证(Fleiss kappa = 0.624,达到实质性一致性);(3) 四个专注风险规避、投资、支出和信用管理的领域智能体,对行为轨迹进行会话级推理,并输出各能力维度的贝叶斯知识追踪估计值;(4) 专家评判智能体整合各领域估计值,生成总体掌握度评分。在193名K-12学生共264个游戏会话中评估,该系统预测的学习进步相关性(r=0.276,p=0.0001)和后测成绩相关性(r=0.333,p<0.0001)显著,且与前测无关,具备收敛效度与区分效度。相比单个LLM基线(r=0.095,不显著),多智能体方法使预测效度提升约三倍,表明领域分解与会话级推理对捕捉财务素养多维特性至关重要。
原文摘要 · Abstract (English)
Assessing financial literacy during gameplay without disrupting the learning experience remains a key challenge in serious games for education. We present the Agentic BKT pipeline, a multi-agent large language model architecture for stealth assessment of financial competencies from open-ended gameplay events. The pipeline processes events from a 2D platformer serious game aligned with the OECD/INFE financial literacy framework through four phases: (1) the game captures every player decision as a structured event log; (2) an LLM event classifier labels each action on a four-point rubric validated against three domain experts (Fleiss kappa = 0.624, substantial agreement); (3) four domain-specific agents specializing in risk mitigation, investing, spending, and credit management perform session-level reasoning over behavioral trajectories, feeding per-competency Bayesian Knowledge Tracing that estimates mastery within each domain; and (4) an expert judge agent synthesizes the domain-level estimates into an overall mastery score. Evaluated with 193 K-12 participants across 264 game sessions, the Agentic BKT pipeline yields mastery estimates significantly correlated with learning gain (r = 0.276, p = 0.0001) and post-test scores (r = 0.333, p < 0.0001) while showing no correlation with pre-test scores, providing both convergent and discriminant validity. The multi-agent approach approximately triples the predictive validity of a single-LLM baseline (r = 0.095, not significant) in this study, demonstrating that domain decomposition and session-level reasoning play a central role in capturing the multidimensional nature of financial literacy from gameplay
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。