提出SMART模型,解决多视图聚类中视图未对齐数据的语义匹配难题
SMART: Semantic Matching Contrastive Learning for Partially View-Aligned Clustering
- 通过语义匹配对比学习缓解跨视图分布偏移
- 在8个基准数据集上均优于现有方法
- 适合处理真实场景中视图不对齐的聚类任务
多视图聚类通过利用数据各视图间的互补信息,已被实证能提升学习性能。然而在现实场景中,严格对齐的视图难以获取,从对齐与非对齐数据中共同学习更具实用性。部分视图对齐聚类(PVC)旨在建立错位视图样本间的对应关系,以更好挖掘视图间的一致性与互补性,涵盖对齐与非对齐数据。但多数现有PVC方法未能有效利用非对齐数据来捕捉同一簇样本间的共享语义。此外,多视图数据的固有异质性导致表示分布偏移,进而影响跨视图潜在特征间有意义对应关系的建立,损害学习效果。为此,我们提出语义匹配对比学习模型(SMART)用于PVC。其核心思想是减轻跨视图分布偏移的影响,从而促进语义匹配对比学习,充分挖掘对齐与非对齐数据中的语义关系。在8个基准数据集上的大量实验表明,本方法在PVC问题上持续优于现有方法。
原文摘要 · Abstract (English)
Multi-view clustering has been empirically shown to improve learning performance by leveraging the inherent complementary information across multiple views of data. However, in real-world scenarios, collecting strictly aligned views is challenging, and learning from both aligned and unaligned data becomes a more practical solution. Partially View-aligned Clustering aims to learn correspondences between misaligned view samples to better exploit the potential consistency and complementarity across views, including both aligned and unaligned data. However, most existing PVC methods fail to leverage unaligned data to capture the shared semantics among samples from the same cluster. Moreover, the inherent heterogeneity of multi-view data induces distributional shifts in representations, leading to inaccuracies in establishing meaningful correspondences between cross-view latent features and, consequently, impairing learning effectiveness. To address these challenges, we propose a Semantic MAtching contRasTive learning model (SMART) for PVC. The main idea of our approach is to alleviate the influence of cross-view distributional shifts, thereby facilitating semantic matching contrastive learning to fully exploit semantic relationships in both aligned and unaligned data. Extensive experiments on eight benchmark datasets demonstrate that our method consistently outperforms existing approaches on the PVC problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。