arXiv:2503.05796cs.CYcs.AI2025-03中稿 · publication, after…被引 1

通过众包研究不同利益相关者对招聘模型评估指标的偏好。

Towards Multi-Stakeholder Evaluation of ML Models: A Crowdsourcing Study on Metric Preferences in Job-matching System

  • 用众包方式让837人对比虚拟招聘模型,评估7个指标的效用。
  • 发现参与者可分5类,不同群体对公平性和性能指标偏好差异显著。
  • 提醒开发者:选评估指标前应征求各方意见,避免片面判断。

机器学习影响多方利益相关者,但缺乏通用评估指标来衡量输出质量(包括性能与公平性)。若仅依赖预设指标而不征求利益相关者意见,会忽视其诉求,导致评估不公。本研究通过众包方式,让837名参与者在20次虚拟招聘系统场景中,从两个假设模型中选择更优者,并计算其对7个指标的效用值。基于效用值将参与者分为5个聚类,分析各群组在指标偏好与共性特征上的差异。结果表明,不同群体对公平性、匹配度等指标的重视程度各异。研究强调,多利益相关方评估模型时,应结合各方意见选择适配指标。

原文摘要 · Abstract (English)

While machine learning (ML) technology affects diverse stakeholders, there is no one-size-fits-all metric to evaluate the quality of outputs, including performance and fairness. Using predetermined metrics without soliciting stakeholder opinions is problematic because it leads to an unfair disregard for stakeholders in the ML pipeline. In this study, to establish practical ways to incorporate diverse stakeholder opinions into the selection of metrics for ML, we investigate participants' preferences for different metrics by using crowdsourcing. We ask 837 participants to choose a better model from two hypothetical ML models in a hypothetical job-matching system twenty times and calculate their utility values for seven metrics. To examine the participants' feedback in detail, we divide them into five clusters based on their utility values and analyze the tendencies of each cluster, including their preferences for metrics and common attributes. Based on the results, we discuss the points that should be considered when selecting appropriate metrics and evaluating ML models with multiple stakeholders.

评估指标利益相关者众包招聘系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。