arXiv:2409.12195stat.MLcs.LG2024-09

复现并验证了基于拓扑结构的高维特征选择算法IVFS,性能优于主流方法。

Reproduction of IVFS algorithm for high-dimensional topology preservation feature selection

  • 通过随机子集思想保持数据拓扑结构以保留相似性
  • 在多数数据集上超越SPEC和MCFS算法
  • 适合关注高维特征选择稳定性的研究者

特征选择是处理高维数据的关键技术。在无监督场景中,许多流行算法聚焦于保持原始数据结构。本文复现了2020年AAAI提出的IVFS算法,该算法受随机子集方法启发,通过维持数据的拓扑结构来保留数据相似性。我们系统整理了IVFS的数学基础,并通过与原论文相似的数值实验验证其有效性。结果表明,IVFS在多数数据集上表现优于SPEC和MCFS,但其收敛性和稳定性问题仍存在。

原文摘要 · Abstract (English)

Feature selection is a crucial technique for handling high-dimensional data. In unsupervised scenarios, many popular algorithms focus on preserving the original data structure. In this paper, we reproduce the IVFS algorithm introduced in AAAI 2020, which is inspired by the random subset method and preserves data similarity by maintaining topological structure. We systematically organize the mathematical foundations of IVFS and validate its effectiveness through numerical experiments similar to those in the original paper. The results demonstrate that IVFS outperforms SPEC and MCFS on most datasets, although issues with its convergence and stability persist.

特征选择拓扑结构高维数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。