让无监督特征选择学会识别因果关系,选对关键特征。
Causally-Aware Unsupervised Feature Selection Learning
- 引入因果正则化重加权样本,平衡混杂分布
- 通过分层聚类区分因果与非因果特征,提升选型准确性
- 融合多粒度相似图,增强模型可解释性,适合高维数据探索
无监督特征选择(UFS)在处理高维无标签数据方面展现出高效性,但现有方法忽视数据中内在的因果机制,导致选取无关特征且可解释性差。此外,基于图的方法未区分因果与非因果特征对相似性图构建的影响,造成虚假连接。为此,本文提出一种新型无监督特征选择方法——因果感知无监督特征选择学习(CAUSE-FS)。CAUSE-FS引入新颖的因果正则化项,对样本进行重加权以平衡每种处理特征的混杂分布,并将其融入广义无监督谱回归模型,从而缓解特征与聚类标签间的虚假关联,实现因果特征选择。同时,采用因果引导的分层聚类,将具有不同因果贡献的特征划分为多粒度层次。通过自适应学习不同粒度下的相似图并融合,强化因果特征在最终融合相似图中的重要性,以捕捉数据可靠的局部结构。大量实验表明,CAUSE-FS在性能上优于现有先进方法,特征可视化进一步验证其可解释性。
原文摘要 · Abstract (English)
Unsupervised feature selection (UFS) has recently gained attention for its effectiveness in processing unlabeled high-dimensional data. However, existing methods overlook the intrinsic causal mechanisms within the data, resulting in the selection of irrelevant features and poor interpretability. Additionally, previous graph-based methods fail to account for the differing impacts of non-causal and causal features in constructing the similarity graph, which leads to false links in the generated graph. To address these issues, a novel UFS method, called Causally-Aware UnSupErvised Feature Selection learning (CAUSE-FS), is proposed. CAUSE-FS introduces a novel causal regularizer that reweights samples to balance the confounding distribution of each treatment feature. This regularizer is subsequently integrated into a generalized unsupervised spectral regression model to mitigate spurious associations between features and clustering labels, thus achieving causal feature selection. Furthermore, CAUSE-FS employs causality-guided hierarchical clustering to partition features with varying causal contributions into multiple granularities. By integrating similarity graphs learned adaptively at different granularities, CAUSE-FS increases the importance of causal features when constructing the fused similarity graph to capture the reliable local structure of data. Extensive experimental results demonstrate the superiority of CAUSE-FS over state-of-the-art methods, with its interpretability further validated through feature visualization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。