轻量化模型+蒸馏量化,让图像压缩在FPGA上高效运行
Lightweight Embedded FPGA Deployment of Learned Image Compression with Knowledge Distillation and Hybrid Quantization
- 用知识蒸馏和单参数调优,快速适配不同硬件平台
- 新激活函数设计,量化后仍保持高编码效率
- 流水线FPGA架构,充分利用资源并提升吞吐
可学习图像压缩(LIC)在率失真效率上已超越标准视频编码器,推动了面向硬件的实现研究。现有LIC硬件实现多侧重延迟与率失真效率平衡,需大量硬件设计探索。本文提出一种新范式:将针对特定平台的调优负担转移到模型尺寸设计,不牺牲率失真效率。首先,构建从参考教师模型蒸馏更轻量学生模型的框架,仅通过调节单一超参数即可满足不同硬件约束,无需复杂设计探索。其次,提出一种面向硬件的广义除法归一化(GDN)激活函数实现,在参数量化后仍保持率失真效率。第三,设计流水线化FPGA配置,通过并行处理与资源优化分配,充分挖掘FPGA资源潜力。实验基于先进LIC模型显示,本方法优于所有现有FPGA实现,且性能接近原始模型。
原文摘要 · Abstract (English)
Learnable Image Compression (LIC) has shown the potential to outperform standardized video codecs in RD efficiency, prompting the research for hardware-friendly implementations. Most existing LIC hardware implementations prioritize latency to RD-efficiency and through an extensive exploration of the hardware design space. We present a novel design paradigm where the burden of tuning the design for a specific hardware platform is shifted towards model dimensioning and without compromising on RD-efficiency. First, we design a framework for distilling a leaner student LIC model from a reference teacher: by tuning a single model hyperparameters, we can meet the constraints of different hardware platforms without a complex hardware design exploration. Second, we propose a hardware-friendly implementation of the Generalized Divisive Normalization - GDN activation that preserves RD efficiency even post parameter quantization. Third, we design a pipelined FPGA configuration which takes full advantage of available FPGA resources by leveraging parallel processing and optimizing resource allocation. Our experiments with a state of the art LIC model show that we outperform all existing FPGA implementations while performing very close to the original model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。