arXiv:2501.13943cs.CLcs.AI2025-01KDD被引 8

用文本描述实现跨领域零样本认知诊断,无需重新训练

Language Representation Favored Zero-Shot Cross-Domain Cognitive Diagnosis

  • 用文本描述学生、题目和知识点,映射到统一语言空间
  • 在多个真实数据集上零样本跨域表现优异,部分媲美全量训练模型
  • 可揭示文理学科、学段间差异,适合教育系统迁移应用

认知诊断旨在根据学生的历史答题记录推断其知识掌握程度。现有认知诊断模型(CDMs)依赖于ID嵌入,通常需针对特定领域单独训练,限制了其在不同学科(如数学、英语、物理)或不同教育平台(如ASSISTments、Junyi Academy、Khan Academy)中的直接应用。本文提出语言表示驱动的零样本跨域认知诊断方法(LRCD)。LRCD首先分析不同领域中学生、题目与概念的行为模式,并用文本描述其特征;通过先进的文本嵌入模块,将这些描述转换为统一语言空间中的向量。为缓解语言空间与认知空间的差异,我们设计语言-认知映射器,学习从前者到后者的转换。由此,这些特征可便捷高效地融入现有CDMs进行联合训练。大量实验表明,基于真实数据集训练的LRCD在多个目标领域实现出色零样本性能,某些情况下甚至可媲美在目标领域全量数据上训练的经典CDMs。值得注意的是,我们意外发现LRCD还能揭示不同学科(如人文与理工)及来源(如小学与中学教育)间的深层差异。

原文摘要 · Abstract (English)

Cognitive diagnosis aims to infer students' mastery levels based on their historical response logs. However, existing cognitive diagnosis models (CDMs), which rely on ID embeddings, often have to train specific models on specific domains. This limitation may hinder their directly practical application in various target domains, such as different subjects (e.g., Math, English and Physics) or different education platforms (e.g., ASSISTments, Junyi Academy and Khan Academy). To address this issue, this paper proposes the language representation favored zero-shot cross-domain cognitive diagnosis (LRCD). Specifically, LRCD first analyzes the behavior patterns of students, exercises and concepts in different domains, and then describes the profiles of students, exercises and concepts using textual descriptions. Via recent advanced text-embedding modules, these profiles can be transformed to vectors in the unified language space. Moreover, to address the discrepancy between the language space and the cognitive diagnosis space, we propose language-cognitive mappers in LRCD to learn the mapping from the former to the latter. Then, these profiles can be easily and efficiently integrated and trained with existing CDMs. Extensive experiments show that training LRCD on real-world datasets can achieve commendable zero-shot performance across different target domains, and in some cases, it can even achieve competitive performance with some classic CDMs trained on the full response data on target domains. Notably, we surprisingly find that LRCD can also provide interesting insights into the differences between various subjects (such as humanities and sciences) and sources (such as primary and secondary education).

认知诊断零样本学习跨域迁移语言建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。