arXiv:2510.04093cs.AI2025-10被引 2

用扩散模型提升大模型在嘈杂教育数据中的认知诊断能力

Harnessing LLM for Noise-Robust Cognitive Diagnosis in Web-Based Intelligent Education Systems

  • 构建双子图融合结构与语义表示,通过两阶段去噪消除噪声干扰
  • 在三个公开教育平台数据集上,不同噪声水平下均实现最优预测性能
  • 适合研究教育智能系统、大模型噪声鲁棒性或认知诊断的学者

基于网络的智能教育系统(WIES)的认知诊断旨在从异构、嘈杂的学生交互数据中评估其对知识概念的掌握程度。尽管已有研究尝试利用大语言模型(LLM)进行认知诊断,但LLM在处理结构化数据时表现不佳,且易受噪声影响导致误判。WIES开放环境持续引入新学生并生成海量响应日志,加剧了传统教育系统固有的数据不平衡和噪声问题。为此,本文提出一种基于扩散的LLM框架DLLM,以实现噪声鲁棒的认知诊断。DLLM首先根据答题正确性构建独立子图,通过关系增强对齐模块缓解数据不平衡;随后将两个子图表示与LLM生成的语义增强表示进行融合与对齐。关键在于,每次对齐前,DLLM采用两阶段去噪扩散模块:先通过无条件扩散去除错误信息,再通过图引导的条件扩散消除误导信息。最终,融合语义知识与结构信息的抗噪表示输入现有诊断模型进行预测。在三个公开教育平台数据集上的实验表明,DLLM在不同噪声水平下均取得最优预测性能,验证了其在保留语义知识的同时具备强噪声鲁棒性。

原文摘要 · Abstract (English)

Cognitive diagnostics in the Web-based Intelligent Education System (WIES) aims to assess students' mastery of knowledge concepts from heterogeneous, noisy interactions. Recent work has tried to utilize Large Language Models (LLMs) for cognitive diagnosis, yet LLMs struggle with structured data and are prone to noise-induced misjudgments. Specially, WIES's open environment continuously attracts new students and produces vast amounts of response logs, exacerbating the data imbalance and noise issues inherent in traditional educational systems. To address these challenges, we propose DLLM, a Diffusion-based LLM framework for noise-robust cognitive diagnosis. DLLM first constructs independent subgraphs based on response correctness, then applies relation augmentation alignment module to mitigate data imbalance. The two subgraph representations are then fused and aligned with LLM-derived, semantically augmented representations. Importantly, before each alignment step, DLLM employs a two-stage denoising diffusion module to eliminate intrinsic noise while assisting structural representation alignment. Specifically, unconditional denoising diffusion first removes erroneous information, followed by conditional denoising diffusion based on graph-guided to eliminate misleading information. Finally, the noise-robust representation that integrates semantic knowledge and structural information is fed into existing cognitive diagnosis models for prediction. Experimental results on three publicly available web-based educational platform datasets demonstrate that our DLLM achieves optimal predictive performance across varying noise levels, which demonstrates that DLLM achieves noise robustness while effectively leveraging semantic knowledge from LLM.

认知诊断大模型去噪教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。