提出可部署的鲁棒特征选择方法,兼顾多群体表现与公平性。
Population-Robust Feature Selection via Generalized Welfare Optimization

- 基于可调福利目标优化,平衡整体性能与弱势群体保护
- 在8个数据集上实现平均与最差群体性能双优
- 适用于大规模特征筛选,支持可解释的公共卫生决策
特征选择是部署决策:同一问卷、检测面板或传感器组合需服务多个异构群体。标准方法通常针对单一主要群体优化,而现有鲁棒方法多为各群体共享统一模型。本文提出PopFS,学习一个共享的可部署特征集,使每个群体可训练专属模型。PopFS采用可调福利目标,在整体预测收益与弱势群体保护间提供灵活权衡。为实现大规模实用,先通过多任务稀疏学习缩减候选池,再通过排序候选增补与替换并仅对短名单全量重训练。在来自五个表格与公共健康数据集的六个预测任务中,八种群体划分下,PopFS持续取得优异平均与最差群体表现,且可扩展至数千候选特征。43州新冠预测研究进一步表明,调整福利目标可显著改善服务最少的州表现,对平均性能影响极小,并带来可解释的症状信号变化。代码已开源。
原文摘要 · Abstract (English)
Choosing which features to collect is a deployment decision: the same limited questionnaire, test panel, or sensor set may need to serve several heterogeneous populations. Standard feature-selection methods typically optimize for one large population, while existing robust approaches tend to learn one shared model for every population. We introduce PopFS, a method for learning one shared, deployable feature set that is robust to population differences while letting each pop- ulation train its own model. PopFS uses a tunable welfare objective that lets practitioners balance overall predictive ben- efit against stronger protection of the populations that benefit least. To make this objective practical at scale, PopFS first uses multitask sparse learning to reduce the candidate pool, then searches directly over hard feature sets by ranking promising additions and swaps and fully refitting only a shortlist. Across eight population splits from six prediction tasks drawn from five tabular and public-health datasets, PopFS consistently achieves strong average and worst-population performance while scaling to thousands of candidate features. A 43-state COVID-19 nowcasting study further shows that changing the welfare objective can improve the least-served states with lit- tle change in average performance and yields an interpretable change in the selected symptom signals. Our code is available at https://github.com/Rachel-Lyu/PopFS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。