arXiv:2502.03220cs.CL2025-02被引 3

统一编码器降低跨语言招聘文本偏见,提升匹配精度。

Mitigating Language Bias in Cross-Lingual Job Retrieval: A Recruitment Platform Perspective

  • 用多任务双编码框架统一学习简历与职位描述的多个文本成分。
  • 模型更小却超越现有最优方法,跨语言性能显著提升。
  • 提出新指标LBKL,有效检测并减少语言偏见,适合招聘系统优化者。

理解简历和职位信息中的文本内容对提升在线招聘平台的岗位匹配准确率和优化搜索系统至关重要。然而,现有方法主要聚焦于单独分析这些信息的各个组成部分,需要多种专用工具处理不同方面,导致方法割裂,可能影响招聘文本处理的整体泛化能力。为此,我们提出一种统一的句子编码器,采用多任务双编码框架,联合学习多个文本成分。实验结果表明,尽管模型规模更小,本方法仍优于其他先进模型。此外,我们提出了一个新的评估指标——语言偏见Kullback-Leibler散度(LBKL),用于衡量编码器中的语言偏见,验证了其在显著降低偏见和提升跨语言性能方面的有效性。

原文摘要 · Abstract (English)

Understanding the textual components of resumes and job postings is critical for improving job-matching accuracy and optimizing job search systems in online recruitment platforms. However, existing works primarily focus on analyzing individual components within this information, requiring multiple specialized tools to analyze each aspect. Such disjointed methods could potentially hinder overall generalizability in recruitment-related text processing. Therefore, we propose a unified sentence encoder that utilized multi-task dual-encoder framework for jointly learning multiple component into the unified sentence encoder. The results show that our method outperforms other state-of-the-art models, despite its smaller model size. Moreover, we propose a novel metric, Language Bias Kullback-Leibler Divergence (LBKL), to evaluate language bias in the encoder, demonstrating significant bias reduction and superior cross-lingual performance.

跨语言招聘系统语言偏见句子编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。