arXiv:2501.07139cs.AIcs.PF2025-01被引 6

FlexQuant让大模型在边缘设备上更省内存、更灵活部署。

FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices

  • 通过生成量化模型集合实现内存弹性,支持动态资源调配。
  • 相比现有方法,内存切换粒度提升15倍,存储成本降低10倍。
  • 兼容主流量化方法,适合资源受限的边缘设备部署。

在边缘设备上部署大语言模型面临严峻挑战。对于共享内存的边缘设备而言,内存弹性至关重要,因为内存资源会动态波动。现有方案或过渡粒度差,或存储开销高。我们提出 FlexQuant,一种新型弹性框架,可生成一组量化模型,相比最先进方法,实现15倍的粒度提升和10倍的存储减少。该框架兼容大多数量化方法,通过剪枝技术在不同存储限制下构建性能-空间权衡选项,显著提升大模型在边缘端的部署灵活性与效率。

原文摘要 · Abstract (English)

Deploying LLMs on edge devices presents serious technical challenges. Memory elasticity is crucial for edge devices with unified memory, where memory is shared and fluctuates dynamically. Existing solutions suffer from either poor transition granularity or high storage costs. We propose FlexQuant, a novel elasticity framework that generates an ensemble of quantized models, providing an elastic hosting solution with 15x granularity improvement and 10x storage reduction compared to SoTA methods. FlexQuant works with most quantization methods and creates a family of trade-off options under various storage limits through our pruning method. It brings great performance and flexibility to the edge deployment of LLMs.

边缘计算量化大模型部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。