发现主流出行预测模型存在种族差异,提出轻量级采样方法减少不公平。
Mind the Gaps: Auditing and Reducing Group Inequity in Large-Scale Mobility Prediction
- 基于人口普查数据构建代理种族标签,实现群体感知采样。
- 在早期采样阶段将群体间性能差距降低最多40%,精度损失极小。
- 适合关注算法公平性的低资源场景研究者与应用开发者。
出行位置预测支撑着越来越多的交通、零售和公共卫生应用,但其社会影响仍缺乏深入探讨。本文审计了基于大规模数据集训练的先进出行预测模型,揭示了用户人口统计特征带来的隐性差异。结合汇总人口普查数据,我们计算了不同族裔群体的预测性能差异,发现底层数据集导致系统性偏差,使不同地理位置和用户群体间的准确率差异显著。为此,我们提出公平引导增量采样(FGIS)策略,适用于增量数据收集场景。由于个体层面的人口统计标签不可用,我们引入尺寸感知K均值(SAKM),在隐式出行空间中聚类用户并强制符合人口普查的群体比例,生成四个主要群体(亚裔、黑人、西班牙裔、白人)的代理种族标签。基于这些标签,采样算法优先选择预期性能提升大且当前代表性不足的用户。该方法逐步构建训练数据集,有效缩小群体间性能差距,同时保持整体准确性。在元路径2向量模型和变压器编码器模型上评估,最多可减少40%的总体不平等,且在早期采样阶段改善最为明显,凸显公平性导向策略在低资源环境中的潜力。研究揭示了出行预测流程中的结构性不公,并证明轻量级、以数据为中心的干预措施可在几乎无额外复杂度下提升公平性,尤其适用于低数据应用场景。
原文摘要 · Abstract (English)
Next location prediction underpins a growing number of mobility, retail, and public-health applications, yet its societal impacts remain largely unexplored. In this paper, we audit state-of-the-art mobility prediction models trained on a large-scale dataset, highlighting hidden disparities based on user demographics. Drawing from aggregate census data, we compute the difference in predictive performance on racial and ethnic user groups and show a systematic disparity resulting from the underlying dataset, resulting in large differences in accuracy based on location and user groups. To address this, we propose Fairness-Guided Incremental Sampling (FGIS), a group-aware sampling strategy designed for incremental data collection settings. Because individual-level demographic labels are unavailable, we introduce Size-Aware K-Means (SAKM), a clustering method that partitions users in latent mobility space while enforcing census-derived group proportions. This yields proxy racial labels for the four largest groups in the state: Asian, Black, Hispanic, and White. Built on these labels, our sampling algorithm prioritizes users based on expected performance gains and current group representation. This method incrementally constructs training datasets that reduce demographic performance gaps while preserving overall accuracy. Our method reduces total disparity between groups by up to 40\% with minimal accuracy trade-offs, as evaluated on a state-of-art MetaPath2Vec model and a transformer-encoder model. Improvements are most significant in early sampling stages, highlighting the potential for fairness-aware strategies to deliver meaningful gains even in low-resource settings. Our findings expose structural inequities in mobility prediction pipelines and demonstrate how lightweight, data-centric interventions can improve fairness with little added complexity, especially for low-data applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。