arXiv:2603.27476cs.AIcs.LG2026-03

首个面向AI人搜平台的多场景评测基准,用真实验证提升评估可信度。

PeopleSearchBench: Evaluating AI-Powered People Search Platforms with Criteria-Grounded Verification

  • 将查询拆解为可独立验证的标准,通过实时网页搜索判断结果真伪
  • 多源搜索代理在网红发现任务中表现显著优于单域系统
  • 评测结果稳定,适合招聘、销售等实际场景的工具选型参考

AI驱动的人力搜索平台在招聘、销售和职业网络中日益普及,但缺乏标准化评估基准。我们提出PeopleSearchBench,一个开源基准,包含119个跨语言查询,覆盖企业招聘、B2B销售、专家查找和影响力人物发现四类场景。核心贡献是基于标准的验证方法(Criteria-Grounded Verification),将每条查询分解为可独立验证的标准,通过实时网络搜索对返回人员进行事实性验证,生成客观相关性判断(与人工标注者的一致性Cohen's kappa=0.84)。我们从相关性精度、有效覆盖率和信息效用三个维度评估了四个架构各异的平台,发现多源搜索代理显著优于单域系统,尤其在影响者发现任务中差距最大。平台排名在评分阈值、维度权重和评判模型的消融实验中均保持稳健。所有代码、查询和评测提示均已公开。

原文摘要 · Abstract (English)

AI-powered people search platforms are increasingly deployed for recruiting, sales prospecting, and professional networking, yet no standardized benchmark exists for their rigorous evaluation. We present PeopleSearchBench, an open-source benchmark comprising 119 multilingual queries across four scenarios: corporate recruiting, B2B sales prospecting, expert search, and influencer discovery. A central contribution is Criteria-Grounded Verification, an evaluation methodology that decomposes each query into explicit, independently checkable criteria and verifies each returned individual via live web search, producing factual relevance judgments rather than subjective LLM-as-judge scores (Cohen's kappa = 0.84 with human annotators). We evaluate four architecturally diverse platforms along three complementary dimensions---Relevance Precision, Effective Coverage, and Information Utility---and find that multi-source search agents significantly outperform single-domain systems, particularly in influencer discovery where the performance gap is largest. Platform rankings are robust across ablations on scoring thresholds, dimension weights, and judge models. All code, queries, and evaluation prompts are publicly available.

人搜评测多源搜索真实验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。