分析13个公平性API在真实项目中的使用与挑战,揭示开发者困境。
Applications and Challenges of Fairness APIs in Machine Learning Software
- 调研204个仓库,梳理13个公平性API的实际应用场景
- 发现开发者普遍缺乏公平性知识,常遇调试难题
- 为教育和工具开发提供实证支持,适合关注AI伦理的研究者
机器学习系统广泛应用于日常生活,尤其在敏感场景中作出影响深远的决策。因此,确保AI/ML系统不产生歧视性结果至关重要。为此,各类开源公平性检测与缓解工具(即公平性API)被开发并使用。本文通过定性研究,探讨这些开源公平性API在实际项目中的应用情境、使用方式及开发者面临的挑战。我们分析了从1885个候选仓库中筛选出的204个使用了13个公平性API的仓库,发现这些工具主要用于学习和解决真实世界问题,覆盖17种具体用例。研究显示,开发者对公平性机制理解不足,频繁遇到调试困难,并常寻求外部建议与资源。研究结果可为未来偏见相关的软件工程研究提供参考,亦有助于指导教育者设计更先进的课程体系。
原文摘要 · Abstract (English)
Machine Learning software systems are frequently used in our day-to-day lives. Some of these systems are used in various sensitive environments to make life-changing decisions. Therefore, it is crucial to ensure that these AI/ML systems do not make any discriminatory decisions for any specific groups or populations. In that vein, different bias detection and mitigation open-source software libraries (aka API libraries) are being developed and used. In this paper, we conduct a qualitative study to understand in what scenarios these open-source fairness APIs are used in the wild, how they are used, and what challenges the developers of these APIs face while developing and adopting these libraries. We have analyzed 204 GitHub repositories (from a list of 1885 candidate repositories) which used 13 APIs that are developed to address bias in ML software. We found that these APIs are used for two primary purposes (i.e., learning and solving real-world problems), targeting 17 unique use-cases. Our study suggests that developers are not well-versed in bias detection and mitigation; they face lots of troubleshooting issues, and frequently ask for opinions and resources. Our findings can be instrumental for future bias-related software engineering research, and for guiding educators in developing more state-of-the-art curricula.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。