构建阿拉伯语招聘文本库,揭示社交媒体招聘的语言规律。
JobArabi: An Arabic Corpus and Analysis of Job Announcements from Social Media
- 基于21类语言关键词从社交平台抓取2万+招聘帖
- 发现性别化用语持续存在、职业需求区域差异明显
- 适合研究阿拉伯语NLP与数字劳动的学者使用
本文介绍JobArabi,一个从2024年1月至2025年10月间收集的阿拉伯语招聘公告大规模语料库,涵盖来自X平台的20,528条公开帖子,覆盖超过两年的阿拉伯语在线就业讨论。该数据集采用语言学驱动的查询框架,涵盖21类反映性别化、复数、正式及方言表达的阿拉伯语招聘关键词。数据包含机构、商业及个人账号发布的帖子,并附带时间戳、互动指标及地理信息(如可用),支持就业话语的时间与地域分析。定量分析揭示了在线招聘中的若干社会语言学模式,包括性别化招聘语言的持续存在、职业需求的区域差异以及招聘内容的情感表达方式。这些发现凸显阿拉伯语社交媒体在研究劳动力市场沟通与语言变迁方面的潜力。JobArabi语料库连同文档与采集脚本将公开发布,以支持阿拉伯语NLP、计算社会科学和数字劳动研究。
原文摘要 · Abstract (English)
This paper introduces JobArabi, a large-scale corpus of Arabic job announcements collected from social media between January 2024 and October 2025. The dataset contains 20,528 public posts from X and captures more than two years of employment-related discourse across Arabic-speaking online communities. The corpus was compiled using a linguistically informed query framework covering 21 Arabic keyword families that reflect gendered, plural, formal, and dialectal expressions of recruitment language. The resulting dataset includes posts from institutional, commercial, and individual accounts and provides metadata such as timestamps, engagement indicators, and geolocation when available, enabling temporal and regional analysis of employment discourse. Quantitative analysis reveals several sociolinguistic patterns in online recruitment, including the persistence of gendered hiring language, regional variation in occupational demand, and the emotional framing of recruitment messages. These findings highlight the potential of Arabic social media as a resource for studying labor market communication and linguistic change. The JobArabi corpus, together with documentation and collection scripts, will be released to support research in Arabic NLP, computational social science, and digital labor studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。