arXiv:2605.21029cs.CL2026-05中稿 · CustomNLP4U 2026

用招聘信息构建AI技能体系,少而精的数据更有效

Building a Custom Taxonomy of AI Skills and Tasks from the Ground Up with Job Postings

论文配图:Building a Custom Taxonomy of AI Skills and Tasks from the Ground Up with Job Postings
图 1 · 摘自论文原文
  • 基于岗位数据设计过滤策略,提升分类清晰度
  • 筛选后数据使领域覆盖更精准,优于原始全部数据
  • 适合想系统化构建职业能力体系的研究者与企业

利用大模型自动化构建分类体系具有广阔前景,但在处理海量快速更新的文本时,如何高效利用数据仍不明确。本文以职场AI技能体系构建为例,使用两大数据规模的岗位招聘文本,研究输入数据筛选对分类体系生成的影响。提出TaxonomyBuilder作为系统性研究框架,评估定制化、数据驱动及分层分类体系的不同配置。结果表明:减少输入数据量反而能提升领域覆盖率——通过筛选输入数据,比直接使用未经筛选的数据给聚类和大模型增强的分层标注工具,带来更清晰的领域映射效果。

原文摘要 · Abstract (English)

Utilizing LLMs for automated taxonomy construction presents a clear opportunity for the comprehensive, yet efficient mapping of potentially complex domains. When contending with high volumes of rapidly growing corpora, however, it becomes unclear how to best leverage such data for optimal taxonomy construction. Taking the case of systematizing AI skills in the workplace, we use two large-scale job postings corpora to investigate key design decisions for the inclusion (or exclusion) of data points for taxonomy construction. We propose TaxonomyBuilder as a blueprint for our systematic study, with which we evaluate various configurations of custom, data-informed, and hierarchical taxonomies. We demonstrate that less data can provide more clarity: filtering inputs to TaxonomyBuilder provides better domain-specific coverage than offering unfiltered inputs to clustering and LLM-enhanced hierarchical taxonomy labeling tools.

技能体系大模型应用数据筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。