arXiv:2608.28652cs.AI2026-08

提出通用优化框架,让低算力设备也能高效运行大模型

A Generalized Optimization Engine (GOE) for Edge AI Inference Acceleration

  • 设计了软硬件无关的通用优化架构,整合多种压缩技术
  • 压缩后语言模型可在无GPU的边缘CPU上运行且保持准确率
  • 强调压缩方法选择比位宽更重要,适合边缘部署场景

人工智能模型在多个领域表现出色,但其广泛部署受限于高昂的计算成本,尤其在资源受限设备上。本文探讨了各类AI模型优化技术、算法和抽象的理论基础,分析其降低计算复杂度、内存占用、延迟和功耗的潜力。进一步提出一种软硬件无关的通用优化架构(GOE),集成多种优化技术以提升效率。研究表明,此类通用优化系统对在战术环境下资源受限的异构硬件上部署模型至关重要。作为具体验证,展示经GOE压缩的语言模型可在无GPU的边缘CPU上部署运行,且任务准确率能否保留取决于压缩方法的选择,而不仅是名义位宽。

原文摘要 · Abstract (English)

Artificial intelligence (AI) models have demonstrated remarkable capabilities across various domains, yet their widespread deployment is impeded by significant computational costs, particularly on resource-constrained devices. This paper explores the theoretical underpinnings of various AI model optimization techniques, algorithms, and abstractions, discussing their potential to reduce computational complexity, memory footprint, latency, and power consumption. Furthermore, we propose a comprehensive hardware (HW) and model-agnostic generalized optimization architecture that integrates these techniques for improved efficiency. Our study underscores the critical role of such a generalized optimization system in preparing model deployment over resource-constrained heterogeneous hardware in a tactical environment. As a concrete demonstration, we show that GOE-compressed language models deploy and run on a GPU-less edge CPU, and that the choice of compression method, not merely its nominal bit-width, determines whether task accuracy survives deployment.

边缘计算模型压缩AI部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。