提出可控制误判率的特征选择统计方法,提升域适应下特征筛选可靠性。
Statistical Inference for Feature Selection after Optimal Transport-based Domain Adaptation
- 基于选择性推断框架,结合线性与二次不等式建模域适应特征选择过程。
- 在显著性水平α=0.05下严格控制假阳性率,同时提升真阳性检测率。
- 适用于小样本目标域场景,适合关注特征选择可信度的研究者。
在域适应(DA)下的特征选择(FS)是机器学习中的关键任务,尤其在目标数据有限时。然而现有方法无法保证FS结果的可靠性。本文提出一种新型统计方法SFS-DA(统计特征选择-域适应),可在预设显著性水平α(如0.05)下严格控制假阳性率(FPR),同时最大化真阳性率。针对域适应对统计推断的影响,SFS-DA通过选择性推断(SI)框架解决,其特征选择过程被建模为线性与二次不等式约束,理论上证明了FPR控制的可行性。进一步引入更优策略以提高检测率。在合成与真实数据集上的实验验证了理论结果,表明SFS-DA性能显著优于现有方法。
原文摘要 · Abstract (English)
Feature Selection (FS) under domain adaptation (DA) is a critical task in machine learning, especially when dealing with limited target data. However, existing methods lack the capability to guarantee the reliability of FS under DA. In this paper, we introduce a novel statistical method to statistically test FS reliability under DA, named SFS-DA (statistical FS-DA). The key strength of SFS-DA lies in its ability to control the false positive rate (FPR) below a pre-specified level $α$ (e.g., 0.05) while maximizing the true positive rate. Compared to the literature on statistical FS, SFS-DA presents a unique challenge in addressing the effect of DA to ensure the validity of the inference on FS results. We overcome this challenge by leveraging the Selective Inference (SI) framework. Specifically, by carefully examining the FS process under DA whose operations can be characterized by linear and quadratic inequalities, we prove that achieving FPR control in SFS-DA is indeed possible. Furthermore, we enhance the true detection rate by introducing a more strategic approach. Experiments conducted on both synthetic and real-world datasets robustly support our theoretical results, showcasing the superior performance of the proposed SFS-DA method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。