对比传统学习,探讨图数据在协作学习中的新方法与挑战
From Euclidean to Graph-Structured Data: A Survey of Collaborative Learning

- 从欧氏数据扩展到图结构数据的协作学习框架
- 提出图数据分布式场景的分类与标准化问题定义
- 适合关注隐私保护与分布式学习的研究者
传统的机器学习方法——集中收集数据、本地训练和推理——面临可扩展性和隐私等根本性限制。为应对这些问题,近年来研究聚焦于联邦学习和去中心化学习等协作学习范式,即各参与方在本地进行训练和推理,仅有限协作。现有研究主要针对具有规则网格结构的欧氏数据(如图像、文本),但这类方法难以捕捉真实世界中大量存在的关系型模式,而这些模式更适合作图结构表示。图学习依赖消息传递机制在连接节点间传播信息,这使其天然契合需要信息交换的协作环境。然而,图结构数据在协作学习中的机遇与挑战仍鲜有研究。本综述系统梳理了从欧氏数据到图结构数据的协作学习演进,首先回顾欧氏数据的基础原则,从学习有效性、效率和隐私保护三方面组织内容;随后拓展至图结构数据,提出图分布场景的分类体系,刻画统计异质性,并建立标准化问题形式与算法框架;最后系统识别开放挑战与未来方向。
原文摘要 · Abstract (English)
The conventional approach to machine learning, that is, collecting data, training models, and performing inference in a single location, faces fundamental limitations, including scalability and privacy, that restrict its applicability. To address these challenges, recent research has explored collaborative learning approaches, including federated learning and decentralized learning, where individual agents perform training and inference locally, with limited collaboration. Most collaborative learning research focuses on Euclidean data with regular, grid-like structure (e.g., images, text). However, these approaches fail to capture the relational patterns in many real-world applications, best represented by graphs. Learning on graphs relies on message-passing mechanisms to propagate information between connected nodes, making it conceptually well-suited for collaborative environments where agents must exchange information. Yet, the opportunities and challenges of learning on graph-structured data in collaborative settings remain largely underexplored. This survey provides a comprehensive investigation of collaborative learning from Euclidean to graph-structured data, aiming to consolidate this emerging field. We begin by reviewing its foundational principles for Euclidean data, organizing them along three core dimensions: learning effectiveness, efficiency, and privacy preservation. We then extend the discussion to graph-structured data, introducing a taxonomy of graph distribution scenarios, characterizing associated statistical heterogeneities, and developing standardized problem formulations and algorithmic frameworks. Finally, we systematically identify open challenges and promising research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。