arXiv:2512.10327cs.CVcs.MM2025-12

提出轻量级选择性填补方法,有效提升不完整多视图聚类性能。

Simple Yet Effective Selective Imputation for Incomplete Multi-view Clustering

  • 先评估再填补:训练无关地判断缺失值是否值得填补
  • 在多个数据集上优于现有填补与非填补方法,尤其在缺失不均衡时表现突出
  • 可直接嵌入现有框架,适合处理缺失数据的多视图学习任务

不完整多视图聚类(IMC)是多视图学习中的关键挑战。主流方法依赖数据填补,但盲目填补可能导致不可靠内容。近期研究采用后填补评估策略:先填补部分或全部缺失值,再通过聚类任务评估质量。该策略计算开销大且依赖聚类模型性能。为此,本文首次提出前评估机制,设计隐式信息量选择性填补(SI$^3$)方法,明确权衡填补效用与风险。SI$^3$以无训练方式评估每个缺失位置的信息量,仅当支持充分时才进行填补。在多视图生成假设下,进一步将选择性填补融入变分推断框架,实现潜在分布层面的不确定性感知填补与鲁棒融合。相比现有方法,SI$^3$轻量、数据驱动、模型无关,可作为即插即用模块集成至现有框架。大量实验表明,其在多个基准数据集上持续超越基于填补和非填补的方法,尤其在缺失不均衡场景中优势显著。

原文摘要 · Abstract (English)

Incomplete Multi-view Clustering (IMC) has emerged as a significant challenge in multi-view learning. A predominant line for IMC is data imputation; however, indiscriminate imputation can result in unreliable content. Recently, researchers have proposed selective imputation methods that use a post-imputation assessment strategy: (1) impute all or some missing values, and (2) evaluate their quality through clustering tasks. We observe that this strategy incurs substantial computational complexity and is heavily dependent on the performance of the clustering model. To address these challenges, we first introduce the concept of pre-imputation assessment. We propose an Implicit Informativeness-based Selective Imputation (SI$^3$) method for incomplete multi-view clustering, which explicitly addresses the trade-off between imputation utility and imputation risk. SI$^3$ evaluates the imputation-relevant informativeness of each missing position in a training-free manner, and selectively imputes data only when sufficient informative support is available. Under a multi-view generative assumption, SI$^3$ further integrates selective imputation into a variational inference framework, enabling uncertainty-aware imputation at the latent distribution level and robust multi-view fusion. Compared with existing selective imputation strategies, SI$^3$ is lightweight, data-driven, and model-agnostic, and can be seamlessly incorporated into existing incomplete multi-view clustering frameworks as a plug-in strategy. Extensive experiments on multiple benchmark datasets demonstrate that SI$^3$ consistently outperforms both imputation-based and imputation-free methods, particularly under challenging unbalanced missing scenarios.

多视图聚类数据填补选择性方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。