arXiv:2603.05262cs.CL2026-03

首个越南岗位招聘数据集,含4.8万条岗位信息,支持AI招聘研究。

VietJobs: A Vietnamese Job Advertisement Dataset

  • 构建覆盖全国34个省市的4.8万条越南岗位数据,含薪资、技能等结构化信息
  • 基于该数据集,大模型在分类与薪资预测任务中表现良好,尤其指令微调模型
  • 为越南语NLP和劳动力市场分析提供重要基准,适合招聘与社会经济研究者

VietJobs是首个大规模公开的越南岗位广告语料库,涵盖48,092条招聘信息和超过1500万字,数据来自越南全部34个省和直辖市。该数据集包含职位名称、类别、薪资、技能要求及雇佣条件等丰富语言与结构化信息,覆盖16个职业领域和全职、兼职、实习等多种就业类型。旨在支持自然语言处理与劳动力市场分析研究,体现显著的语言、区域与社会经济多样性。我们在两个核心任务上对多个生成式大模型(LLMs)进行基准测试:职位类别分类与薪资估算。指令微调模型如Qwen2.5-7B-Instruct和Llama-SEA-LION-v3-8B-IT在少样本和微调设置下表现优异,但暴露出多语言及越南语特有建模在结构化劳动力市场预测中的挑战。VietJobs为越南语NLP建立了新基准,为未来招聘语言、社会经济表征及人工智能驱动的劳动力市场分析研究提供了宝贵基础。所有代码与资源均开放于:https://github.com/VinNLP/VietJobs。

原文摘要 · Abstract (English)

VietJobs is the first large-scale, publicly available corpus of Vietnamese job advertisements, comprising 48,092 postings and over 15 million words collected from all 34 provinces and municipalities across Vietnam. The dataset provides extensive linguistic and structured information, including job titles, categories, salaries, skills, and employment conditions, covering 16 occupational domains and multiple employment types (full-time, part-time, and internship). Designed to support research in natural language processing and labour market analytics, VietJobs captures substantial linguistic, regional, and socio-economic diversity. We benchmark several generative large language models (LLMs) on two core tasks: job category classification and salary estimation. Instruction-tuned models such as Qwen2.5-7B-Instruct and Llama-SEA-LION-v3-8B-IT demonstrate notable gains under few-shot and fine-tuned settings, while highlighting challenges in multilingual and Vietnamese-specific modelling for structured labour market prediction. VietJobs establishes a new benchmark for Vietnamese NLP and offers a valuable foundation for future research on recruitment language, socio-economic representation, and AI-driven labour market analysis. All code and resources are available at: https://github.com/VinNLP/VietJobs.

越南语NLP岗位数据集劳动力市场大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。