arXiv:2509.13763cs.LG2025-09

从因果视角出发,解决多视图无监督特征选择中的伪相关问题。

Beyond Correlation: Causal Multi-View Unsupervised Feature Selection Learning

  • 引入因果模型识别并分离混淆因子,避免误选无关特征。
  • 通过共识聚类与因果正则化联合优化,提升特征选择可靠性。
  • 首次在无监督多视图场景下系统研究因果特征选择,适合数据降维研究者。

多视图无监督特征选择(MUFS)近年来备受关注,因其在多视图无标签数据上实现降维的潜力。现有方法通常依赖特征与聚类标签间的相关性来选择判别性特征,但一个关键而未被充分探讨的问题是:这些相关性是否足够可靠以指导特征选择?本文从因果视角分析MUFS,提出一种新型结构因果模型,揭示现有方法可能因忽略由混淆因子引起的伪相关而选择无关特征。基于此,我们提出新方法CAUSA(CAusal multi-view Unsupervised feature Selection leArning)。首先,采用广义无监督谱回归模型,通过捕捉特征与共识聚类标签之间的依赖关系识别信息特征;其次,引入因果正则化模块,自适应分离多视图数据中的混淆因子,并学习共享样本权重以平衡混淆因子分布,从而缓解伪相关。最后,将两者整合进统一学习框架,实现因果信息特征的选择。大量实验表明,CAUSA优于多个前沿方法。据我们所知,这是首个在无监督设置下深入研究多视图因果特征选择的工作。

原文摘要 · Abstract (English)

Multi-view unsupervised feature selection (MUFS) has recently received increasing attention for its promising ability in dimensionality reduction on multi-view unlabeled data. Existing MUFS methods typically select discriminative features by capturing correlations between features and clustering labels. However, an important yet underexplored question remains: \textit{Are such correlations sufficiently reliable to guide feature selection?} In this paper, we analyze MUFS from a causal perspective by introducing a novel structural causal model, which reveals that existing methods may select irrelevant features because they overlook spurious correlations caused by confounders. Building on this causal perspective, we propose a novel MUFS method called CAusal multi-view Unsupervised feature Selection leArning (CAUSA). Specifically, we first employ a generalized unsupervised spectral regression model that identifies informative features by capturing dependencies between features and consensus clustering labels. We then introduce a causal regularization module that can adaptively separate confounders from multi-view data and simultaneously learn view-shared sample weights to balance confounder distributions, thereby mitigating spurious correlations. Thereafter, integrating both into a unified learning framework enables CAUSA to select causally informative features. Comprehensive experiments demonstrate that CAUSA outperforms several state-of-the-art methods. To our knowledge, this is the first in-depth study of causal multi-view feature selection in the unsupervised setting.

多视图学习因果推断特征选择无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。