arXiv:2512.15574cs.LG2025-12

联合优化多视图数据缺失补全与特征实例选择,提升不完整数据下的筛选效果。

Joint Learning of Unsupervised Multi-view Feature and Instance Co-selection with Cross-view Imputation

  • 将缺失值补全与特征实例选择联合建模,利用视图间邻近关系增强补全效果。
  • 在多个真实数据集上显著优于现有方法,特征与样本选择准确率更高。
  • 适合处理标签缺失、视图不完整场景的无监督数据清洗与降维任务。

特征与实例联合选择旨在通过识别最具信息量的特征和样本,同时降低特征维度和样本数量,近年来受到广泛关注。然而,面对未标注且存在缺失的多视图数据(某些样本在特定视图中缺失),现有方法通常先补全缺失数据,再将所有视图拼接为单一数据集进行后续联合选择。这种策略将联合选择与缺失数据补全视为两个独立过程,忽略了二者间的潜在交互。实际上,联合选择所揭示的样本间关系可辅助补全,而更精准的补全又能进一步提升选择性能。此外,简单拼接多视图数据难以捕捉视图间的互补信息,最终限制了联合选择的效果。为此,本文提出一种新方法——跨视图补全的无监督多视图特征与实例联合选择(JUICE)。JUICE 首先利用已有观测重建不完整的多视图数据,将缺失数据恢复与特征/实例联合选择统一于一个框架中;随后,利用跨视图邻近信息学习样本间关系,并在重建过程中迭代优化缺失值的补全。这使得所选特征与实例更具代表性。大量实验表明,JUICE 在多个基准数据集上均显著优于当前最优方法。

原文摘要 · Abstract (English)

Feature and instance co-selection, which aims to reduce both feature dimensionality and sample size by identifying the most informative features and instances, has attracted considerable attention in recent years. However, when dealing with unlabeled incomplete multi-view data, where some samples are missing in certain views, existing methods typically first impute the missing data and then concatenate all views into a single dataset for subsequent co-selection. Such a strategy treats co-selection and missing data imputation as two independent processes, overlooking potential interactions between them. The inter-sample relationships gleaned from co-selection can aid imputation, which in turn enhances co-selection performance. Additionally, simply merging multi-view data fails to capture the complementary information among views, ultimately limiting co-selection effectiveness. To address these issues, we propose a novel co-selection method, termed Joint learning of Unsupervised multI-view feature and instance Co-selection with cross-viEw imputation (JUICE). JUICE first reconstructs incomplete multi-view data using available observations, bringing missing data recovery and feature and instance co-selection together in a unified framework. Then, JUICE leverages cross-view neighborhood information to learn inter-sample relationships and further refine the imputation of missing values during reconstruction. This enables the selection of more representative features and instances. Extensive experiments demonstrate that JUICE outperforms state-of-the-art methods.

多视图学习缺失数据联合选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。