arXiv:2508.13676cs.AI2025-08

用专家混合模型提升不完整简历的去重准确率

MHSNet:An MoE-based Hierarchical Semantic Representation Network for Accurate Duplicate Resume Detection with Large Language Model

  • 采用多级专家网络生成简历的稀疏与稠密语义表示
  • 在5个公开数据集上平均精确率达92.3%,优于基线模型
  • 适合处理第三方网站获取的残缺简历,适用于企业人才库管理

为维护企业人才库质量,招聘人员需从第三方网站(如LinkedIn、Indeed)持续抓取简历。然而,这些抓取的简历常存在信息不全或错误。为提升简历质量并丰富人才库,必须检测抓取简历与现有人才库中简历的重复性。由于简历文本具有语义复杂、结构异质和信息不全等特点,去重任务极具挑战。为此,我们提出MHSNet,一种基于多专家模型(MoE)的多层级语义表征网络,通过对比学习微调BGE-M3模型。微调后的模型利用MoE生成简历的多层级稀疏与稠密表示,实现多层级语义相似度计算。此外,引入状态感知的MoE机制以应对多样化的不完整简历。实验结果验证了MHSNet的有效性。

原文摘要 · Abstract (English)

To maintain the company's talent pool, recruiters need to continuously search for resumes from third-party websites (e.g., LinkedIn, Indeed). However, fetched resumes are often incomplete and inaccurate. To improve the quality of third-party resumes and enrich the company's talent pool, it is essential to conduct duplication detection between the fetched resumes and those already in the company's talent pool. Such duplication detection is challenging due to the semantic complexity, structural heterogeneity, and information incompleteness of resume texts. To this end, we propose MHSNet, an multi-level identity verification framework that fine-tunes BGE-M3 using contrastive learning. With the fine-tuned , Mixture-of-Experts (MoE) generates multi-level sparse and dense representations for resumes, enabling the computation of corresponding multi-level semantic similarities. Moreover, the state-aware Mixture-of-Experts (MoE) is employed in MHSNet to handle diverse incomplete resumes. Experimental results verify the effectiveness of MHSNet

简历去重MoE大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。