用正弦激活提升量化低秩适配器表达力,显著减少模型传输内存。
SineLoRA$Δ$: Sine-Activated Delta Compression
- 引入正弦激活增强量化后低秩适配器的稳定秩
- 在多领域实现最高66%内存压缩,性能相近
- 首次理论分析量化对稳定秩影响,适合高效模型部署
资源受限的权重部署具有重要实际意义。近期研究关注 extit{Delta Compression}任务,即各方持有共同基础模型,仅需通信压缩后的权重更新。然而,主流参数高效方法如低秩适配(LoRA)存在固有表示限制,尤其在激进量化下更为明显。为此,我们基于近期工作,利用固定频率正弦函数在不增加参数的前提下提升稳定秩。我们将该思想扩展至量化场景,首次给出量化下稳定秩演化的理论分析。据此提出SineLoRA$Δ$,一种原理严谨且高效的增量压缩方法,通过正弦激活提升量化低秩适配器的表达能力。我们在语言建模、视觉-语言任务及文生图等多个领域验证其有效性,实现最高66%的内存缩减,且性能保持相近。此外,我们创新性地应用标准Bjøntegaard Delta指标,在率失真曲线上一致比较适配器压缩效果。
原文摘要 · Abstract (English)
Resource-constrained weight deployment is a task of immense practical importance. Recently, there has been interest in the specific task of \textit{Delta Compression}, where parties each hold a common base model and only communicate compressed weight updates. However, popular parameter efficient updates such as Low Rank Adaptation (LoRA) face inherent representation limitations - which are especially pronounced when combined with aggressive quantization. To overcome this, we build on recent work that improves LoRA representation capacity by using fixed-frequency sinusoidal functions to increase stable rank without adding additional parameters. We extend this to the quantized setting and present the first theoretical analysis showing how stable rank evolves under quantization. From this, we introduce SineLoRA$Δ$, a principled and effective method for delta compression that improves the expressivity of quantized low-rank adapters by applying a sinusoidal activation. We validate SineLoRA$Δ$ across a diverse variety of domains - including language modeling, vision-language tasks, and text-to-image generation - achieving up to 66% memory reduction with similar performance. We additionally provide a novel application of the canonical Bjøntegaard Delta metric to consistently compare adapter compression changes across the rate-distortion curve.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。