arXiv:2605.04761cs.LGcs.AI2026-05被引 1

为学生构建可解释的思维模型,通过人机协作提升个性化教育体验

Cognitive Twins: Investigating Personalized Thinking Model Building and Its Performance Enhancement with Human-in-the-Loop

论文配图:Cognitive Twins: Investigating Personalized Thinking Model Building and Its Performance Enhancement with Human-in-the-Loop
图 1 · 摘自论文原文
  • 基于五层结构整合学习日志,用大模型与聚类技术构建认知孪生体
  • 人机协同后模型准确率提升至75.48%,用户满意度达4.30分(5分制)
  • 模型展现从行为到价值观的语义抽象特征,适合教育智能系统研发者

本文提出个性化思维模型(PTM),一种分层可解释的学习者表征框架,用于支持人工智能教育。PTM将学习者日志中的证据组织为五个层级:行为实例、行为模式、认知习惯、元认知倾向和自我系统价值。该模型基于马扎诺的新教育目标分类体系,旨在克隆学习者的思维模式并构建认知孪生体。通过结合Gemini 2.5 Pro大模型推理、句子嵌入、降维与共识聚类的流程构建。在为期七周、40名参与者的研究中,通过三种方法评估其保真度:首先,自动评估使用原子信息点匹配,未经过人机协同(HITL)前的总体F1得分为74.57%,经优化后升至75.48%;其次,用户评价采用李克特量表,预处理与后处理条件下的平均评分分别为4.26和4.30;第三,语义对齐验证显示,主题连贯性从行为层的0.436上升至核心价值层的0.626,而词汇重叠率则从0.114下降至0.007。结果表明,PTM输出具有可接受保真度,被用户普遍认为反映其真实思维,并呈现出与语义抽象一致的层次化特征。

原文摘要 · Abstract (English)

This paper presents the Personalized Thinking Model (PTM), a hierarchical and interpretable learner representation designed for AI supported education. PTM organizes evidence from learner journals into a five-layer structure covering behavioral instances, behavioral patterns, cognitive routines, metacognitive tendencies, and self-system values. PTM is grounded in Marzano's New Taxonomy of Educational Objectives and tries to clone learner's thinking model and build cognitive twin. It was constructed using a pipeline that combines large language model inference (Gemini 2.5 Pro), sentence embeddings, dimensionality reduction, and consensus clustering. This paper evaluates PTM fidelity through three methods applied to 40 participants in a seven-week study. First, automatic evaluation using atomic information point matching yielded an overall F1 score of 74.57% before human-in-the-loop (HITL) refinement and 75.48% after refinement. Second, user evaluation using a Likert scale produced mean ratings of 4.26 and 4.30 on a five-point scale for pre and post-HITL conditions respectively. Third, semantic alignment verification showed that topic coherence increased from 0.436 at the behavioral layer to 0.626 at the core value layer, while lexical overlap with journal vocabulary decreased from 0.114 to 0.007 across those same layers. These results suggest that the PTM produces outputs with acceptable fidelity, was generally perceived by users as reflecting their thinking, and showed a pattern consistent with semantic abstraction across layers.

个性化教育认知建模人机协同思维模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。