arXiv:2410.05018cs.IRcs.CL2024-10中稿 · the 4th Workshop o…

专家查找系统评估存在偏差,需改进标注方法以确保公平比较。

On the Biased Assessment of Expert Finding Systems

  • 用系统推荐知识领域辅助标注,易导致性能虚高。
  • 传统关键词模型在该标注下表现被夸大,新神经模型对比失效。
  • 添加同义词后发现对字面匹配有强偏好,适合评估者参考。

在大型组织中,识别特定主题的专家对挖掘跨团队与部门的知识资源至关重要。企业级专家检索系统基于员工海量异构数据自动发现并构建其专业能力。评估这些系统需全面的专家标注作为真实标签,但获取困难,因此常依赖系统推荐的知识领域进行验证。本案例研究分析了此类推荐如何影响评估结果:在主流基准上,系统验证的标注使传统基于关键词的检索模型性能被高估,甚至破坏与更先进神经方法的可比性。通过引入同义词扩展知识领域,发现存在对术语字面匹配的显著偏差。我们提出标注过程约束方案,可在避免偏差的同时保留推荐的实用价值。研究提醒应谨慎设计或选择基准,以确保专家查找方法的合理比较。

原文摘要 · Abstract (English)

In large organisations, identifying experts on a given topic is crucial in leveraging the internal knowledge spread across teams and departments. So-called enterprise expert retrieval systems automatically discover and structure employees' expertise based on the vast amount of heterogeneous data available about them and the work they perform. Evaluating these systems requires comprehensive ground truth expert annotations, which are hard to obtain. Therefore, the annotation process typically relies on automated recommendations of knowledge areas to validate. This case study provides an analysis of how these recommendations can impact the evaluation of expert finding systems. We demonstrate on a popular benchmark that system-validated annotations lead to overestimated performance of traditional term-based retrieval models and even invalidate comparisons with more recent neural methods. We also augment knowledge areas with synonyms to uncover a strong bias towards literal mentions of their constituent words. Finally, we propose constraints to the annotation process to prevent these biased evaluations, and show that this still allows annotation suggestions of high utility. These findings should inform benchmark creation or selection for expert finding, to guarantee meaningful comparison of methods.

专家查找评估偏差标注方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。