用语言模型自动优化无线联邦学习中的设备选择,省电又高效。
Tool-Aided Evolutionary LLM for Generative Policy Toward Efficient Resource Management in Wireless Federated Learning
- 用自然语言提示引导大模型生成设备选择策略,跨场景泛化强。
- 实测能耗降低32%,在动态网络中仍保持稳定性能。
- 虚拟环境训练减少真实交互,适合资源受限的边缘部署。
联邦学习(FL)可在隐私保护下实现边缘设备间的分布式模型训练。然而其效率高度依赖于动态异构无线环境中的有效设备选择与高维资源分配。传统方法需领域知识、大量超参数调优或高交互成本。本文提出工具辅助进化大语言模型(T-ELLM)框架,用于生成无线联邦学习环境下的合格设备选择策略。与传统优化方法不同,T-ELLM利用自然语言场景提示提升在不同网络条件下的泛化能力。该框架在数学上解耦联合优化问题,使设备选择策略可有效学习,同时将资源分配委托给凸优化工具。为增强适应性,T-ELLM集成了一种样本高效的基于模型的虚拟学习环境,捕捉设备选择与学习性能间的关系,支持后续群体相对策略优化。此协同方法减少了对真实世界交互的依赖,降低通信开销,同时保持高保真决策。理论分析证明虚拟与真实环境间的差异有界,确保虚拟环境中学习的优势函数与现实条件偏差极小。实验结果表明,T-ELLM在能效方面优于基准方法,并对环境变化表现出强鲁棒性。
原文摘要 · Abstract (English)
Federated Learning (FL) enables distributed model training across edge devices in a privacy-friendly manner. However, its efficiency heavily depends on effective device selection and high-dimensional resource allocation in dynamic and heterogeneous wireless environments. Conventional methods demand a confluence of domain-specific expertise, extensive hyperparameter tuning, and/or heavy interaction cost. This paper proposes a Tool-aided Evolutionary Large Language Model (T-ELLM) framework to generate a qualified policy for device selection in a wireless FL environment. Unlike conventional optimization methods, T-ELLM leverages natural language-based scenario prompts to enhance generalization across varying network conditions. The framework decouples the joint optimization problem mathematically, enabling tractable learning of device selection policies while delegating resource allocation to convex optimization tools. To improve adaptability, T-ELLM integrates a sample-efficient, model-based virtual learning environment that captures the relationship between device selection and learning performance, facilitating subsequent group relative policy optimization. This concerted approach reduces reliance on real-world interactions, minimizing communication overhead while maintaining high-fidelity decision-making. Theoretical analysis proves that the discrepancy between virtual and real environments is bounded, ensuring the advantage function learned in the virtual environment maintains a provably small deviation from real-world conditions. Experimental results demonstrate that T-ELLM outperforms benchmark methods in energy efficiency and exhibits robust adaptability to environmental changes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。