用大模型精准提取课程能力并匹配岗位需求,发现教育与就业的差距
An NLP-Driven Framework for Curriculum-Labor Market Alignment: Schema-Constrained LLM Extraction, ESCO-Anchored Semantic Matching, and Multi-Dimensional Gap Quantification
- 基于结构化提示和双模型协作,从课程文本中提取400条能力条目
- 识别出通用技能缺口达25.0%,算法理论缺13.8%,软件工程缺12.2%
- 可为高校课程改革提供量化依据,适合教育评估与人才规划者使用
从多样化的教育与劳动力市场语料中进行受结构约束的信息抽取仍是自然语言处理中的开放挑战。现有方法多依赖表层词汇匹配,难以捕捉隐含能力,缺乏共享分类体系支撑,且无法衡量抽取可靠性或文档完整度。本文提出四阶段NLP框架:(i)利用双模型前沿大模型联合对齐JSON Schema约束的七维度能力形式化;(ii)通过Sentence-BERT(SBERT)将提取记录与11个领域的ESCO v1.2.1受控词表对齐;(iii)采用两级仲裁机制解决模型分歧;(iv)结合逐槽位Cohen's kappa、结构符合性与文档完整性审计进行验证。该框架应用于阿联酋大学计算机科学本科项目(ABET认证)的课程-就业对齐评估。从85门课程的2025-2026教学计划中提取400条能力记录,并在五个分析层次(从核心计算到概率加权学生路径)下,与483个条款的30份职位招聘启事在SBERT余弦阈值0.50下对齐。提取器在技能槽上达到0.79的Cohen's kappa,100%结构符合率,100%文档完整度。结果显示通用与跨领域能力存在25.0%供需缺口,算法与计算理论缺13.8%,软件工程与项目管理缺12.2%,而人工智能与数据科学虽仅覆盖38.6%供给,但缺口仅为1.8%。
原文摘要 · Abstract (English)
Schema-constrained information extraction from diverse educational and labor-market corpora remains an open challenge in natural language processing because existing pipelines rely primarily on lexical-surface methods that cannot recover implicit competencies, lack grounding in shared taxonomies, and provide no formal measures of extraction reliability or document-level completeness. To address these limitations, this paper proposes a four-stage NLP framework that combines (i) schema-constrained prompting of a two-model frontier-LLM ensemble against a JSON Schema-enforced seven-slot competency formalism, (ii) Sentence-BERT (SBERT) alignment of the extracted records against an eleven-domain ESCO v1.2.1 controlled vocabulary, (iii) a two-tier adjudication protocol that resolves inter-model disagreements, and (iv) a verification mechanism that combines per-slot Cohen's kappa, schema conformance, and document-level completeness audits. The framework is instantiated for a critical application in higher-education quality assurance, namely curriculum-labor market alignment for the ABET-accredited BSc Computer Science program at the United Arab Emirates University. The pipeline extracts 400 competency records from the 85-course 2025-2026 study plan and aligns them, under a five-scope analysis ranging from the computing core to a probability-weighted student trajectory, with 30 job postings (483 requirement clauses) at an SBERT cosine threshold of 0.50. The extractor achieves Cohen's kappa of 0.79 on the skill slot, with 100% schema conformance and 100% document-level completeness. The alignment surfaces interpretable supply-demand gaps of 25.0% in general and transversal skills, 13.8% in algorithms and computational theory, and 12.2% in software engineering and project management, with a near-zero 1.8% gap in artificial intelligence and data science despite 38.6% supply coverage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。