提出安全鲁棒特征选择方法,确保模型在环境变化时仍保留关键特征。
Safe Distributionally Robust Feature Selection under Covariate Shift
- 基于分布鲁棒优化,构建可抵御输入分布偏移的特征筛选框架。
- 在有限样本下保证不误剔重要特征,理论证明无假排除风险。
- 适合工业多传感器系统,应对部署环境与训练环境不一致场景。
在实际机器学习中,模型开发与部署环境常存在差异,尤其当模型被众多用户在多样环境中使用时。保持模型在可能部署环境中的可靠性能,称为分布鲁棒(DR)学习。本文研究分布鲁棒特征选择(DRFS),聚焦于由工业需求驱动的稀疏传感应用。在实际多传感器系统中,通常基于大量可用传感器的性能评估,在部署前选定共享传感器子集。部署时,各用户可能根据特定环境对模型进行适应或微调。当部署环境偏离开发阶段预期时,该策略可能导致系统缺少最优性能所需的关键传感器。为此,我们提出安全-分布鲁棒特征选择(safe-DRFS),将安全筛选从传统稀疏建模扩展至协变量偏移下的分布鲁棒设置。该方法识别出一个特征子集,涵盖在指定输入分布偏移范围内可能成为最优的所有子集,并提供有限样本下的理论保证:不会错误剔除任何关键特征。
原文摘要 · Abstract (English)
In practical machine learning, the environments encountered during the model development and deployment phases often differ, especially when a model is used by many users in diverse settings. Learning models that maintain reliable performance across plausible deployment environments is known as distributionally robust (DR) learning. In this work, we study the problem of distributionally robust feature selection (DRFS), with a particular focus on sparse sensing applications motivated by industrial needs. In practical multi-sensor systems, a shared subset of sensors is typically selected prior to deployment based on performance evaluations using many available sensors. At deployment, individual users may further adapt or fine-tune models to their specific environments. When deployment environments differ from those anticipated during development, this strategy can result in systems lacking sensors required for optimal performance. To address this issue, we propose safe-DRFS, a novel approach that extends safe screening from conventional sparse modeling settings to a DR setting under covariate shift. Our method identifies a feature subset that encompasses all subsets that may become optimal across a specified range of input distribution shifts, with finite-sample theoretical guarantees of no false feature elimination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。