构建多语言简历与职位描述问答基准,评估大模型在招聘场景的阅读理解能力。
JobResQA: A Benchmark for LLM Machine Reading Comprehension on Multilingual Résumés and JDs
- 基于真实数据合成多语言简历与职位描述对,生成581个跨语言问答对。
- 英语和西班牙语表现较好,其他语言性能显著下降,暴露多语言理解短板。
- 支持公平性研究,适合关注招聘系统公平性与多语言AI的开发者使用。
我们提出JobResQA,一个用于评估大模型在人力资源场景下阅读理解能力的多语言问答基准,涵盖简历与职位描述。数据集包含105对合成简历-职位描述,覆盖英文、西班牙语、意大利语、德语和中文五种语言,共581个问答对,问题涵盖从基础事实提取到复杂跨文档推理三个难度层级。通过去标识化与数据合成构建真实且保护隐私的数据,采用占位符控制人口统计与职业属性,便于系统性偏见与公平性研究。提出基于TEaR方法的低成本人机协作翻译流程,结合MQM错误标注与选择性后编辑,确保高质量多语言平行数据。使用大模型作为裁判者对多个开源大模型进行基线评估,结果显示英语和西班牙语表现更优,其余语言性能显著下降,揭示了当前多语言大模型在人力资源应用中的关键差距。该基准已公开:https://github.com/Avature/jobresqa-benchmark。
原文摘要 · Abstract (English)
We introduce JobResQA, a multilingual Question Answering benchmark for evaluating Machine Reading Comprehension (MRC) capabilities of LLMs on HR-specific tasks involving résumés and job descriptions. The dataset comprises 581 QA pairs across 105 synthetic résumé-job description pairs in five languages (English, Spanish, Italian, German, and Chinese), with questions spanning three complexity levels from basic factual extraction to complex cross-document reasoning. We propose a data generation pipeline derived from real-world sources through de-identification and data synthesis to ensure both realism and privacy, while controlled demographic and professional attributes (implemented via placeholders) enable systematic bias and fairness studies. We also present a cost-effective, human-in-the-loop translation pipeline based on the TEaR methodology, incorporating MQM error annotations and selective post-editing to ensure an high-quality multi-way parallel benchmark. We provide a baseline evaluations across multiple open-weight LLM families using an LLM-as-judge approach revealing higher performances on English and Spanish but substantial degradation for other languages, highlighting critical gaps in multilingual MRC capabilities for HR applications. JobResQA provides a reproducible benchmark for advancing fair and reliable LLM-based HR systems. The benchmark is publicly available at: https://github.com/Avature/jobresqa-benchmark
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。