构建阿拉伯语求职广告多国数据集,助力公平性与劳动力研究
ArabJobs: A Multinational Corpus of Arabic Job Ads
- 收集埃及、约旦、沙特、阿联酋四国超8500条招聘广告
- 揭示性别角色与职业分布差异,展现方言多样性特征
- 支持薪资预测、职业分类等下游任务,适合语言与社会研究者
ArabJobs 是一个公开的阿拉伯语求职广告语料库,涵盖埃及、约旦、沙特阿拉伯和阿联酋,包含超过8,500条招聘信息,总词数逾55万。该数据集反映了阿拉伯劳动力市场的语言、地域与社会经济差异。本文分析了性别代表性与职业结构,并突出广告中方言的变异,为未来研究提供契机。我们还展示了利用大语言模型进行薪资估算与职业类别归一化的应用,并设立了性别偏见检测与职业分类的基准任务。结果表明,ArabJobs 对公平性敏感的阿拉伯语自然语言处理及劳动力市场研究具有重要价值。数据集已开源:https://github.com/drelhaj/ArabJobs。
原文摘要 · Abstract (English)
ArabJobs is a publicly available corpus of Arabic job advertisements collected from Egypt, Jordan, Saudi Arabia, and the United Arab Emirates. Comprising over 8,500 postings and more than 550,000 words, the dataset captures linguistic, regional, and socio-economic variation in the Arab labour market. We present analyses of gender representation and occupational structure, and highlight dialectal variation across ads, which offers opportunities for future research. We also demonstrate applications such as salary estimation and job category normalisation using large language models, alongside benchmark tasks for gender bias detection and profession classification. The findings show the utility of ArabJobs for fairness-aware Arabic NLP and labour market research. The dataset is publicly available on GitHub: https://github.com/drelhaj/ArabJobs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。