审计高校风险模型发现:年轻男性和国际生被错误标记为高风险
Fairness Audits of Institutional Risk Models in Deployed ML Pipelines

- 用机构数据复现预警系统,全程追踪公平性偏差
- 年轻男性、国际生被过度标记,而同风险女性与年长生被忽略
- 后处理环节将概率压缩成等级,放大了不公平结果
对部署中的早期预警系统(EWS)进行基于复制的公平性审计,使用中心学院多年合作积累的机构训练数据与设计规范重建模型。通过标准公平性指标,在整个流程(训练数据、模型预测、后处理)中评估性别、年龄、居住状态带来的差异。审计发现系统存在系统性误配:年轻、男性及国际学生被过度标记为需支持对象,即使其中许多人最终成功;而年龄较大、女性学生在相同退学风险下却未被充分识别。后处理阶段通过将异质概率合并为百分位风险层级,加剧了不公平。本研究提供可复现的审计方法,揭示偏差如何在各阶段累积,强调需同时评估构念效度与统计公平性。该成果是探索高等教育中算法、学生数据与权力关系的系列研究之一。
原文摘要 · Abstract (English)
Fairness audits of institutional risk models are critical for understanding how deployed machine learning pipelines allocate resources. Drawing on multi-year collaboration with Centennial College, where our prior ethnographic work introduced the ASP-HEI Cycle, we present a replica-based audit of a deployed Early Warning System (EWS), replicating its model using institutional training data and design specifications. We evaluate disparities by gender, age, and residency status across the full pipeline (training data, model predictions, and post-processing) using standard fairness metrics. Our audit reveals systematic misallocation: younger, male, and international students are disproportionately flagged for support, even when many ultimately succeed, while older and female students with comparable dropout risk are under-identified. Post-processing amplifies these disparities by collapsing heterogeneous probabilities into percentile-based risk tiers. This work provides a replicable methodology for auditing institutional ML systems and shows how disparities emerge and compound across stages, highlighting the importance of evaluating construct validity alongside statistical fairness. It contributes one empirical thread to a broader program investigating algorithms, student data, and power in higher education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。