提出可控制误报率的序列特征选择检验方法,保障结果可靠性。
Statistical Inference for Sequential Feature Selection after Domain Adaptation
- 基于域适应后序列特征选择,设计统计检验方法
- 在α=0.05下严格控制误报率,提升检验功效
- 适用于AIC/BIC/调整R²等模型选择准则,适合高维数据研究者
在高维回归中,序列特征选择(SeqFS)常用于识别相关特征。当数据有限时,域适应(DA)通过迁移相关源域知识至目标域,提升泛化性能。尽管序贯特征选择与域适应结合是机器学习中的重要任务,现有方法无法保证结果的可靠性。本文提出一种针对SeqFS-DA所选特征的新型统计检验方法,其核心优势在于能将误报率(FPR)控制在显著性水平α(如0.05)以下。此外,引入策略以增强检验的统计功效。进一步地,该方法扩展至使用AIC、BIC及调整R²作为模型选择标准的序贯特征选择场景。在合成与真实数据集上进行了大量实验,验证了理论结果并展示了方法的优越性能。
原文摘要 · Abstract (English)
In high-dimensional regression, feature selection methods, such as sequential feature selection (SeqFS), are commonly used to identify relevant features. When data is limited, domain adaptation (DA) becomes crucial for transferring knowledge from a related source domain to a target domain, improving generalization performance. Although SeqFS after DA is an important task in machine learning, none of the existing methods can guarantee the reliability of its results. In this paper, we propose a novel method for testing the features selected by SeqFS-DA. The main advantage of the proposed method is its capability to control the false positive rate (FPR) below a significance level $α$ (e.g., 0.05). Additionally, a strategic approach is introduced to enhance the statistical power of the test. Furthermore, we provide extensions of the proposed method to SeqFS with model selection criteria including AIC, BIC, and adjusted R-squared. Extensive experiments are conducted on both synthetic and real-world datasets to validate the theoretical results and demonstrate the proposed method's superior performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。