系统梳理多视图聚类方法,揭示其优势与挑战。
Advanced Unsupervised Learning: A Comprehensive Overview of Multi-View Clustering Techniques
- 按协同训练、核方法等7类对多视图聚类进行分类
- 分析了140+篇论文,对比早融合、晚融合等集成策略
- 适合关注无监督学习、跨领域数据挖掘的研究者
机器学习面临计算限制、单视图算法局限及多源异构大数据处理复杂等挑战。多视图聚类(MVC)作为一类无监督多视图学习方法,可弥补单视图方法不足,提供更丰富的数据表征,有效解决多种无监督学习任务。相比传统单视图方法,多视图数据语义丰富,实用性强,尽管其本身结构复杂。本综述做出三方面贡献:(1) 将多视图聚类方法系统划分为协同训练、协同正则化、子空间、深度学习、核方法、锚点法和图方法等明确类别;(2) 深入分析各类方法的优势、劣势及实际挑战,如可扩展性与数据不完整问题;(3) 展望新兴趋势、跨学科应用及未来研究方向。研究涵盖超过140篇基础与近期文献,对比早融合、晚融合与联合学习等集成策略,并系统考察医疗、多媒体、社交网络分析等领域的应用实例。通过整合这些工作,旨在填补现有研究空白,为该领域发展提供可操作的洞见。
原文摘要 · Abstract (English)
Machine learning techniques face numerous challenges to achieve optimal performance. These include computational constraints, the limitations of single-view learning algorithms and the complexity of processing large datasets from different domains, sources or views. In this context, multi-view clustering (MVC), a class of unsupervised multi-view learning, emerges as a powerful approach to overcome these challenges. MVC compensates for the shortcomings of single-view methods and provides a richer data representation and effective solutions for a variety of unsupervised learning tasks. In contrast to traditional single-view approaches, the semantically rich nature of multi-view data increases its practical utility despite its inherent complexity. This survey makes a threefold contribution: (1) a systematic categorization of multi-view clustering methods into well-defined groups, including co-training, co-regularization, subspace, deep learning, kernel-based, anchor-based, and graph-based strategies; (2) an in-depth analysis of their respective strengths, weaknesses, and practical challenges, such as scalability and incomplete data; and (3) a forward-looking discussion of emerging trends, interdisciplinary applications, and future directions in MVC research. This study represents an extensive workload, encompassing the review of over 140 foundational and recent publications, the development of comparative insights on integration strategies such as early fusion, late fusion, and joint learning, and the structured investigation of practical use cases in the areas of healthcare, multimedia, and social network analysis. By integrating these efforts, this work aims to fill existing gaps in MVC research and provide actionable insights for the advancement of the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。