从1.55亿份招聘广告中提取100亿条劳动力市场数据,构建可公开使用的月度分析数据集。
Extracting O*NET Features from the NLx Corpus to Build Public Use Aggregate Labor Market Data
- 基于O*NET框架,用NLP从招聘广告中提取结构化信息
- 从1.55亿份广告中提取超100亿条数据,覆盖技能、薪资、工具等
- 开源工具包支持研究者复现,适合教育与职业发展领域应用
在线招聘广告数据难以获取且缺乏标准化和透明性。标准职业信息数据库(O*NET)更新缓慢,依赖小样本调查。本文采用O*NET作为框架,构建自然语言处理工具,从招聘广告中提取结构化信息。我们发布了开源的招聘广告分析工具包(Job Ad Analysis Toolkit, JAAT),并在跨样本和大模型评分测试中验证其可靠性与准确性。利用国家劳动力交易所(NLx)研究枢纽提供的1.55亿份在线招聘广告,成功提取超过100亿条数据点,涵盖O*NET任务、职业代码、工具技术、薪资、技能、行业等特征。本文还构建了2015至2025年按月统计的职位、州和行业层级聚合数据集,可用于教育与劳动力发展研究。
原文摘要 · Abstract (English)
Data from online job postings are difficult to access and are not built in a standard or transparent manner. Data included in the standard taxonomy and occupational information database (O*NET) are updated infrequently and based on small survey samples. We adopt O*NET as a framework for building natural language processing tools that extract structured information from job postings. We publish the Job Ad Analysis Toolkit (JAAT), a collection of open-source tools built for this purpose, and demonstrate its reliability and accuracy in out-of-sample and LLM-as-a-Judge testing. We extract more than 10 billion data points from more than 155 million online job ads provided by the National Labor Exchange (NLx) Research Hub, including O*NET tasks, occupation codes, tools, and technologies, as well as wages, skills, industry, and more features. We describe the construction of a dataset of occupation, state, and industry level features aggregated by monthly active jobs from 2015 - 2025. We illustrate the potential for research and future uses in education and workforce development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。