构建统一平台解决特征选择评估的可复现难题
MH-FSF: A Unified Framework for Overcoming Benchmarking and Reproducibility Limitations in Feature Selection Evaluation
- 提出模块化框架MH-FSF,集成17种特征选择方法
- 在10个公开安卓恶意软件数据集上系统评估性能差异
- 适合安全、数据挖掘领域研究者复现与对比实验
特征选择对构建高效预测模型至关重要,能降低维度并突出关键特征。然而当前研究常受限于基准测试不足和依赖专有数据集,严重阻碍可复现性,并可能影响整体性能。为此,我们提出MH-FSF框架,一个全面、模块化且可扩展的平台,旨在促进特征选择方法的复现与实现。该框架由协作研究开发,包含17种方法(11种经典,6种领域特定),并在10个公开安卓恶意软件数据集上实现系统评估。结果表明,平衡与不平衡数据集间性能存在显著差异,凸显了数据预处理及选择标准需考虑此类不对称性的重要性。本工作证明了统一平台在比较多种特征选择技术中的关键作用,有助于提升方法论的一致性与严谨性。通过提供此框架,我们希望显著拓展现有文献,并为特征选择研究开辟新方向,尤其在安卓恶意软件检测领域。
原文摘要 · Abstract (English)
Feature selection is vital for building effective predictive models, as it reduces dimensionality and emphasizes key features. However, current research often suffers from limited benchmarking and reliance on proprietary datasets. This severely hinders reproducibility and can negatively impact overall performance. To address these limitations, we introduce the MH-FSF framework, a comprehensive, modular, and extensible platform designed to facilitate the reproduction and implementation of feature selection methods. Developed through collaborative research, MH-FSF provides implementations of 17 methods (11 classical, 6 domain-specific) and enables systematic evaluation on 10 publicly available Android malware datasets. Our results reveal performance variations across both balanced and imbalanced datasets, highlighting the critical need for data preprocessing and selection criteria that account for these asymmetries. We demonstrate the importance of a unified platform for comparing diverse feature selection techniques, fostering methodological consistency and rigor. By providing this framework, we aim to significantly broaden the existing literature and pave the way for new research directions in feature selection, particularly within the context of Android malware detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。