用BERT和降维优化职业文本聚类,自动映射不同定义的职业
Improving Clustering on Occupational Text Data through Dimensionality Reduction
- 结合BERT与多种聚类方法构建职业映射管道
- 降维显著提升聚类性能,尤其在Silhouette指标上改进明显
- 适合需要职业自动分类的跨系统迁移场景
本研究针对美国知名职业数据库O*NET中的职业定义,提出一种优化的聚类机制。尽管所有职业均基于严谨的美国调查定义,但在不同公司和国家中其定义可能存在差异。若需扩展已有的O*NET数据以涵盖不同任务定义的职业,建立定义间的映射至关重要。我们提出一个融合多种BERT技术与聚类方法的流程,并评估了不同降维方法对聚类算法性能指标的影响。最终通过引入专用轮廓系数(silhouette)方法进一步提升结果。该基于聚类与降维的职业映射新方法可实现职业的自动化区分,为有职业转换需求的人群开辟新路径。
原文摘要 · Abstract (English)
In this study, we focused on proposing an optimal clustering mechanism for the occupations defined in the well-known US-based occupational database, O*NET. Even though all occupations are defined according to well-conducted surveys in the US, their definitions can vary for different firms and countries. Hence, if one wants to expand the data that is already collected in O*NET for the occupations defined with different tasks, a map between the definitions will be a vital requirement. We proposed a pipeline using several BERT-based techniques with various clustering approaches to obtain such a map. We also examined the effect of dimensionality reduction approaches on several metrics used in measuring performance of clustering algorithms. Finally, we improved our results by using a specialized silhouette approach. This new clustering-based mapping approach with dimensionality reduction may help distinguish the occupations automatically, creating new paths for people wanting to change their careers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。