不同目标函数会显著影响无监督特征选择的搜索效果和结果质量。
Objective-Induced Bias and Search Dynamics in Multiobjective Unsupervised Feature Selection

- 用三种评估目标对比大小正则化方向,研究目标设计的影响。
- 轮廓系数易导致低维平凡解,而PCA损失可得高精度紧凑子集。
- 适合关注无监督特征选择中目标函数设计的研究者参考。
无监督特征选择通常被建模为多目标优化问题,同时优化子集质量和大小。然而该方法的行为高度依赖于评估目标的选择、子集大小正则化的方向以及初始化策略。本文在包含已知信息、冗余和无关特征的合成数据集上,通过六种组合(三种评估目标:准确率、轮廓系数、PCA重构损失,搭配子集大小最小化或最大化)进行控制实验。结果表明,不同公式对搜索动态和最终帕累托前沿质量有显著影响。基于轮廓系数的公式表现出强烈偏向低基数平凡解,且难以反映预测性能;而提出的PCA损失目标能生成紧凑子集,测试准确率与直接优化监督准确率所得子集相当。研究证实目标设计是有效多目标无监督特征选择的核心。
原文摘要 · Abstract (English)
Unsupervised feature selection is commonly formulated as a multiobjective optimisation problem that jointly optimises subset quality and subset size. Yet the behaviour of this formulation depends critically on the choice of evaluation objective, the direction of subset-size regularisation, and the initialisation strategy. We study these factors in a controlled setting using a synthetic dataset with known informative, redundant, and irrelevant feature types. Six formulations are compared by combining three evaluation objectives: accuracy, silhouette score, and PCA reconstruction loss with subset-size minimisation or maximisation. The results show that formulation strongly affects both search dynamics and the quality of the resulting Pareto front. Silhouette-based formulations exhibit a strong bias toward trivial low-cardinality solutions and remain weak proxies for predictive performance. In contrast, the proposed PCA loss objective produces compact subsets with test accuracy comparable to subsets obtained by directly optimising supervised accuracy. Overall, the study shows that objective design is central to effective multiobjective unsupervised feature selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。