提出新型特征选择算法,能同时捕捉基因间的复杂互作关系。
Advancing Interaction-Sensitive Feature Selection: Novel Relief-Based Algorithms, Expanded Comparisons, and Recommendations for Biomedical Data Mining

- 设计多种改进的Relief算法,通过优化邻近样本和评分策略提升性能。
- 在模拟基因数据中验证,新算法对主效应和两两互作均具有强检测能力。
- 适用于高维生物医学数据,帮助构建更高效、可解释的预测模型。
在高维生物医学数据分析中,可靠的特征选择可降低计算成本、提升建模性能,并生成更简洁可解释的模型。然而,多数基于过滤的方法难以检测特征互作,而包裹式或嵌入式方法则计算开销大。Relief基算法(RBAs)作为过滤方法,对特征互作敏感且避免了其他局限。本研究(1)重构并扩展了scikit-rebate Python工具包,集成现有及新提出的RBA变体;(2)在多样化基因组模拟数据上进行严格的基准比较。扩展后包含SWRF*、mu-Relief及5种新RBA变体,采用不同的邻居选择与特征评分策略。所有算法在不同样本量、特征数、遗传力及关联类型(如主效应与互作)的数据集上评估其预测特征排序与运行时间。除mu-Relief外,所有RBA在噪声数据中均能有效检测二阶互作;采用'远距离'评分策略的RBA表现最佳(如MultiSWRFDB*),但对主效应不敏感。SWRF、MultiSWRF、MultiSURF和MultiSWRFDB在主效应与二阶互作数据中表现最优,其中MultiSWRFDB在三阶互作下也领先。scikit-rebate重构使运行时间减少10至35倍。新引入的RBA表现出色,能稳健保留主效应与二阶表观遗传互作信号,为下游建模提供可靠支持。
原文摘要 · Abstract (English)
As a precursor to high-dimensional biomedical data modeling, reliable feature selection can reduce computational expense, improve modeling performance, and yield simpler, more interpretable models. However, most filter-based feature selection methods struggle to detect feature interactions, while wrapper or embedded feature selection methods are computationally expensive. Relief-based algorithms (RBAs) are filter methods that are sensitive to feature interactions while mitigating these other limitations. This study (1) refactors, optimizes, and expands the scikit-rebate Python package with existing and newly proposed RBA variants and (2) conducts rigorous RBA benchmark comparisons across diverse genomic simulations. We expand scikit-rebate to include SWRF*, mu-Relief, and 5 novel RBA variants implementing alternative strategies for neighbor selection and feature scoring. All RBAs were evaluated to compare predictive feature ranking and runtime across simulated genomic datasets varying in sample size, number of features, heritability, and underlying association type (e.g. main effects and interactions). All RBAs, except mu-Relief, were proficient in detecting 2-way interactions in noisy data. RBAs utilizing 'far' scoring were best at detecting 2-way interactions - with MultiSWRFDB* top-performing - but were far less sensitive to main effects. SWRF, MultiSWRF, MultiSURF, and MultiSWRFDB yielded top performance across main effect and 2-way interaction datasets with MultiSWRFDB performing best when also considering 3-way interactions. Refactoring of scikit-rebate resulted in 10 to 35-fold reductions in RBA runtimes. The newly introduced RBAs were among the strongest performing, and by robustly retaining both main effects and 2-way epistatic interactions, these algorithms preserve predictive signals for downstream modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。