针对异构多模态数据,提出按数据量自适应的联邦学习框架,提升训练效率与精度。
Progressive Size-Adaptive Federated Learning: A Comprehensive Framework for Heterogeneous Multi-Modal Data Systems
- 根据数据集大小动态调整训练策略,实现跨模态自适应联邦学习。
- 在1000-1500样本区间表现最优,超2000样本时性能显著下降。
- 适用于医疗、传感器等结构化数据场景,尤其适合资源受限系统。
联邦学习(FL)作为分布式机器学习的变革性范式,在保护数据隐私的同时广泛应用。然而,现有方法主要关注模型异构性与聚合技术,忽视了数据集规模对联邦训练动态的根本影响。本文提出基于数据量自适应的联邦学习(SAFL)框架,系统性地依据异构多模态数据的数据集规模特征组织联邦学习过程。在涵盖7种模态(视觉、文本、时间序列、音频、传感器、医学图像、多模态)的13个不同数据集上进行综合评估发现:1)联邦学习有效性最佳的数据集规模为1000-1500样本;2)结构化数据(时间序列、传感器)性能显著优于非结构化数据(文本、多模态);3)数据集超过2000样本后性能系统性下降。SAFL在所有数据集上平均准确率达87.68%,结构化模态达到99%以上。框架通信效率优越,总数据传输仅7.38 GB,共558次通信,同时保持高性能。实时监控系统揭示了系统资源利用、网络效率与训练动态的深层规律。本工作填补了数据特征驱动联邦学习策略的空白,为神经网络与学习系统在真实场景中的部署提供理论与实践指导。
原文摘要 · Abstract (English)
Federated Learning (FL) has emerged as a transformative paradigm for distributed machine learning while preserving data privacy. However, existing approaches predominantly focus on model heterogeneity and aggregation techniques, largely overlooking the fundamental impact of dataset size characteristics on federated training dynamics. This paper introduces Size-Based Adaptive Federated Learning (SAFL), a novel progressive training framework that systematically organizes federated learning based on dataset size characteristics across heterogeneous multi-modal data. Our comprehensive experimental evaluation across 13 diverse datasets spanning 7 modalities (vision, text, time series, audio, sensor, medical vision, and multimodal) reveals critical insights: 1) an optimal dataset size range of 1000-1500 samples for federated learning effectiveness; 2) a clear modality performance hierarchy with structured data (time series, sensor) significantly outperforming unstructured data (text, multimodal); and 3) systematic performance degradation for large datasets exceeding 2000 samples. SAFL achieves an average accuracy of 87.68% across all datasets, with structured data modalities reaching 99%+ accuracy. The framework demonstrates superior communication efficiency, reducing total data transfer to 7.38 GB across 558 communications while maintaining high performance. Our real-time monitoring framework provides unprecedented insights into system resource utilization, network efficiency, and training dynamics. This work fills critical gaps in understanding how data characteristics should drive federated learning strategies, providing both theoretical insights and practical guidance for real-world FL deployments in neural network and learning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。