arXiv:2505.23801cs.CLcs.AI2025-05被引 1

针对异构NLP任务,提出高效低通信的联邦学习框架

SEMFED: Semantic-Aware Resource-Efficient Federated Learning for Heterogeneous NLP Tasks

  • 根据语义多样性和设备资源动态选客户端
  • 模型自适应压缩,通信量降低80.5%且精度超98%
  • 适合资源受限、数据分布差异大的真实场景

联邦学习(FL)在保护数据隐私的同时训练模型,但应用于自然语言处理(NLP)时面临语义异构、词汇不匹配及边缘设备资源不均等挑战。本文提出SEMFED,一种面向异构NLP任务的语义感知、资源高效的联邦学习框架。其核心创新包括:(1) 融合语义多样性与资源约束的客户端选择机制;(2) 针对设备能力自适应调整的NLP专用模型结构,保留语义信息;(3) 通信高效的语义特征压缩技术,显著降低带宽消耗。在多个NLP分类任务上的实验表明,SEMFED在通信成本降低80.5%的前提下,模型精度保持在98%以上,优于现有先进方法。结果证明,SEMFED能有效应对计算资源、网络稳定性与语义数据分布各异的客户端环境,适用于实际联邦NLP部署。

原文摘要 · Abstract (English)

Background: Federated Learning (FL) has emerged as a promising paradigm for training machine learning models while preserving data privacy. However, applying FL to Natural Language Processing (NLP) tasks presents unique challenges due to semantic heterogeneity across clients, vocabulary mismatches, and varying resource constraints on edge devices. Objectives: This paper introduces SEMFED, a novel semantic-aware resource-efficient federated learning framework specifically designed for heterogeneous NLP tasks. Methods: SEMFED incorporates three key innovations: (1) a semantic-aware client selection mechanism that balances semantic diversity with resource constraints, (2) adaptive NLP-specific model architectures tailored to device capabilities while preserving semantic information, and (3) a communication-efficient semantic feature compression technique that significantly reduces bandwidth requirements. Results: Experimental results on various NLP classification tasks demonstrate that SEMFED achieves an 80.5% reduction in communication costs while maintaining model accuracy above 98%, outperforming state-of-the-art FL approaches. Conclusion: SEMFED effectively manages heterogeneous client environments with varying computational resources, network reliability, and semantic data distributions, making it particularly suitable for real-world federated NLP deployments.

联邦学习NLP资源效率通信压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。