跨语言职位匹配新模型,支持四国语言精准对齐
Multilingual JobBERT for Cross-Lingual Job Title Matching
- 基于对比学习构建多语言职位标题匹配模型
- 在2100万条职位数据上训练,跨语言性能超越基线
- 无需特定任务标注,适合多语言求职市场分析
我们提出JobBERT-V3,一种基于对比学习的跨语言职位标题匹配模型。在领先的单语模型JobBERT-V2基础上,通过合成翻译和超过2100万条职位标题的平衡多语言数据集,扩展支持英语、德语、西班牙语和中文。该模型保留了前代的高效架构,无需任务特定监督即可实现跨语言稳健对齐。在TalentCLEF 2025基准上的广泛评估表明,JobBERT-V3优于强大多语言基线,在单语和跨语言设置中均表现一致。此外,模型还可有效为给定职位标题排序相关技能,展示了其在多语言劳动力市场智能中的更广泛应用前景。模型已公开:https://huggingface.co/TechWolf/JobBERT-v3。
原文摘要 · Abstract (English)
We introduce JobBERT-V3, a contrastive learning-based model for cross-lingual job title matching. Building on the state-of-the-art monolingual JobBERT-V2, our approach extends support to English, German, Spanish, and Chinese by leveraging synthetic translations and a balanced multilingual dataset of over 21 million job titles. The model retains the efficiency-focused architecture of its predecessor while enabling robust alignment across languages without requiring task-specific supervision. Extensive evaluations on the TalentCLEF 2025 benchmark demonstrate that JobBERT-V3 outperforms strong multilingual baselines and achieves consistent performance across both monolingual and cross-lingual settings. While not the primary focus, we also show that the model can be effectively used to rank relevant skills for a given job title, demonstrating its broader applicability in multilingual labor market intelligence. The model is publicly available: https://huggingface.co/TechWolf/JobBERT-v3.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。