arXiv:2602.00135cs.CV2026-02中稿 · ICLR被引 9

用傅里叶变换联合压缩多模态模型,提升精度与效率。

LLaVA-FA: Learning Fourier Approximation for Compressing Large Multimodal Models

  • 在频域中联合低秩分解与量化,减少误差累积。
  • 相比现有方法,在多个基准上性能更优且参数激活极少。
  • 适合需要高效部署大模型的场景,如移动端或多设备推理。

大型多模态模型(LMMs)在视觉-语言任务中表现卓越,但其巨大的计算与内存开销限制了实际应用。现有压缩方法通常将低秩分解与量化解耦,导致重建误差叠加,尤其在具有跨模态冗余的架构中更为显著。为此,我们提出LLaVA-FA,一种在频域中联合执行低秩与量化近似的新型高效多模态模型。利用傅里叶变换的去相关性和共轭对称性,该方法实现更紧凑且精确的权重表示。此外,我们引入PolarQuant——一种针对复数矩阵设计的极坐标量化方法,并提出可选的对角校准(ODC)方案,无需大规模校准数据。大量实验表明,所提方法在多个基准上优于现有高效多模态模型,同时保持极少的激活参数和低计算成本,验证了其作为压缩大型多模态模型的有效解决方案。

原文摘要 · Abstract (English)

Large multimodal models (LMMs) have achieved impressive performance on various vision-language tasks, but their substantial computational and memory costs hinder their practical deployment. Existing compression methods often decouple low-rank decomposition and quantization, leading to compounded reconstruction errors, especially in multimodal architectures with cross-modal redundancy. To address this issue, we propose LLaVA-FA, a novel efficient LMM that performs joint low-rank plus quantization approximation in the frequency domain. By leveraging the de-correlation and conjugate symmetry properties of Fourier transform, LLaVA-FA achieves more compact and accurate weight representations. Furthermore, we introduce PolarQuant, a polar-coordinate quantization method tailored for complex matrices, and an optional diagonal calibration (ODC) scheme that eliminates the need for large-scale calibration data. Extensive experimental results demonstrate that our proposed LLaVA-FA outperforms existing efficient multimodal models across multiple benchmarks while maintaining minimal activated parameters and low computational costs, validating its effectiveness as a powerful solution for compressing LMMs.

模型压缩多模态傅里叶变换量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。