arXiv:2507.11515cs.LGcs.AI2025-07被引 1

用扩散模型动态调整低秩适配参数,远程微调更省带宽。

AirLLM: Diffusion Policy-based Adaptive LoRA for Remote Fine-Tuning of LLM over the Air

  • 基于扩散模型生成自适应的低秩矩阵秩配置
  • 在不同信噪比下传输成本降低,微调效果更优
  • 适合边缘设备远程高效微调大模型的场景

在边缘设备上运行大语言模型面临通信带宽有限、计算与内存资源紧张的挑战,因此云端辅助的远程微调变得不可或缺。然而,现有低秩适配(LoRA)方法通常采用固定或启发式秩配置,且所有LoRA参数的空中传输效率较低。为此,我们提出AirLLM,一种面向通信感知的分层扩散策略框架。具体地,将秩配置建模为涵盖所有LoRA插入投影的结构化动作向量。针对高维序列决策问题,采用近端策略优化(PPO)代理联合观测无线状态与语言复杂度生成粗粒度决策,并通过去噪扩散隐式模型(DDIM)细化为高分辨率、任务与信道自适应的秩向量。两个模块交替优化,其中DDIM在无分类器引导(CFG)框架下训练,以保持与PPO奖励的一致性。在不同信噪比下的实验表明,AirLLM持续提升微调性能的同时显著降低传输开销,验证了强化学习驱动、扩散细化的秩自适应在空对空远程微调中的有效性与可扩展性。

原文摘要 · Abstract (English)

Operating Large Language Models (LLMs) on edge devices is increasingly challenged by limited communication bandwidth and strained computational and memory costs. Thus, cloud-assisted remote fine-tuning becomes indispensable. Nevertheless, existing Low-Rank Adaptation (LoRA) approaches typically employ fixed or heuristic rank configurations, and the subsequent over-the-air transmission of all LoRA parameters could be rather inefficient. To address this limitation, we develop AirLLM, a hierarchical diffusion policy framework for communication-aware LoRA adaptation. Specifically, AirLLM models the rank configuration as a structured action vector that spans all LoRA-inserted projections. To solve the underlying high-dimensional sequential decision-making problem, a Proximal Policy Optimization (PPO) agent generates coarse-grained decisions by jointly observing wireless states and linguistic complexity, which are then refined via Denoising Diffusion Implicit Models (DDIM) to produce high-resolution, task- and channel-adaptive rank vectors. The two modules are optimized alternatively, with the DDIM trained under the Classifier-Free Guidance (CFG) paradigm to maintain alignment with PPO rewards. Experiments under varying signal-to-noise ratios demonstrate that AirLLM consistently enhances fine-tuning performance while significantly reducing transmission costs, highlighting the effectiveness of reinforcement-driven, diffusion-refined rank adaptation for scalable and efficient remote fine-tuning over the air.

边缘计算低秩适配扩散模型远程微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。