大模型无需训练就能准确预测技能前置关系,助力个性化学习系统。
How Well Do LLMs Predict Prerequisite Skills? Zero-Shot Comparison to Expert-Defined Concepts
- 用自然语言描述零样本预测技能依赖关系,不需额外训练。
- LLaMA4-Maverick等模型与专家标注匹配度超70%,推理延迟低。
- 适合教育科技、智能辅导系统开发者快速构建技能图谱。
前置技能——掌握高级概念前所需的基础能力——对有效学习、评估和技能差距分析至关重要。传统上由领域专家定义,但维护成本高且难以扩展。本文研究大语言模型(LLMs)在零样本设置下预测前置技能的可行性,仅使用自然语言描述,无需任务微调。我们构建了基于ESCO分类体系的基准数据集ESCO-PrereqSkill,包含3,196个技能及其专家定义的前置链接。采用标准化提示策略,评估13个先进LLM(如GPT-4、Claude 3、Gemini、LLaMA 4、Qwen2、DeepSeek),涵盖语义相似度、BERTScore及推理延迟。结果表明,LLaMA4-Maverick、Claude-3-7-Sonnet、Qwen2-72B等模型的预测结果与专家真值高度一致,展现强大的无监督语义推理能力。这证明了LLMs在个性化学习、智能辅导和基于技能的推荐系统中实现可扩展技能建模的巨大潜力。
原文摘要 · Abstract (English)
Prerequisite skills - foundational competencies required before mastering more advanced concepts - are important for supporting effective learning, assessment, and skill-gap analysis. Traditionally curated by domain experts, these relationships are costly to maintain and difficult to scale. This paper investigates whether large language models (LLMs) can predict prerequisite skills in a zero-shot setting, using only natural language descriptions and without task-specific fine-tuning. We introduce ESCO-PrereqSkill, a benchmark dataset constructed from the ESCO taxonomy, comprising 3,196 skills and their expert-defined prerequisite links. Using a standardized prompting strategy, we evaluate 13 state-of-the-art LLMs, including GPT-4, Claude 3, Gemini, LLaMA 4, Qwen2, and DeepSeek, across semantic similarity, BERTScore, and inference latency. Our results show that models such as LLaMA4-Maverick, Claude-3-7-Sonnet, and Qwen2-72B generate predictions that closely align with expert ground truth, demonstrating strong semantic reasoning without supervision. These findings highlight the potential of LLMs to support scalable prerequisite skill modeling for applications in personalized learning, intelligent tutoring, and skill-based recommender systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。