通过精选数据与低碳节点,降低联邦学习的碳排放。
Eco-Friendly AI: Unleashing Data Power for Green Federated Learning
- 按数据质量筛选子集,减少训练数据量。
- 实测可降低联邦学习碳排放,提升能效。
- 适合关注绿色AI与可持续计算的研究者。
人工智能与机器学习的广泛应用带来了显著的环境影响,尤其体现在能源消耗和碳排放上。训练数据规模是影响模型训练能耗的关键因素。由于传感器和设备持续生成海量数据,联邦学习(FL)通过不传输原始数据实现模型训练,既降低了通信成本又增强了隐私保护。然而,数据源异构性(如数据量与质量差异)、计算节点能力不均及环境影响仍带来挑战。本文提出一种以数据为中心的绿色联邦学习方法,聚焦于减少训练数据量以降低环境影响。方法包括分析联邦数据特征、基于质量指标选取最优数据子集,并选择环境影响最小的参与节点。构建了综合评估体系,研究数据质量与体积对训练性能及碳排放的影响。在此基础上,开发交互式推荐系统,通过数据缩减优化联邦配置,最小化训练过程中的环境足迹。应用于时间序列分类任务,结果表明该方法有效降低了联邦学习的碳排放。
原文摘要 · Abstract (English)
The widespread adoption of Artificial Intelligence (AI) and Machine Learning (ML) comes with a significant environmental impact, particularly in terms of energy consumption and carbon emissions. This pressing issue highlights the need for innovative solutions to mitigate AI's ecological footprint. One of the key factors influencing the energy consumption of ML model training is the size of the training dataset. ML models are often trained on vast amounts of data continuously generated by sensors and devices distributed across multiple locations. To reduce data transmission costs and enhance privacy, Federated Learning (FL) enables model training without the need to move or share raw data. While FL offers these advantages, it also introduces challenges due to the heterogeneity of data sources (related to volume and quality), computational node capabilities, and environmental impact. This paper contributes to the advancement of Green AI by proposing a data-centric approach to Green Federated Learning. Specifically, we focus on reducing FL's environmental impact by minimizing the volume of training data. Our methodology involves the analysis of the characteristics of federated datasets, the selecting of an optimal subset of data based on quality metrics, and the choice of the federated nodes with the lowest environmental impact. We develop a comprehensive methodology that examines the influence of data-centric factors, such as data quality and volume, on FL training performance and carbon emissions. Building on these insights, we introduce an interactive recommendation system that optimizes FL configurations through data reduction, minimizing environmental impact during training. Applying this methodology to time series classification has demonstrated promising results in reducing the environmental impact of FL tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。