arXiv:2503.20245cs.ARcs.AI2025-03被引 4

8K高清视频实时超分加速器,动态选网节省算力

ESSR: An 8K@30FPS Super-Resolution Accelerator With Edge Selective Network

  • 根据图像边缘特征动态选择子网络,降低计算量
  • 实现8K@30FPS流畅处理,功耗仅0.2075W,能效达4797Mpixels/J
  • 适合移动端和嵌入式设备的高分辨率视频增强场景

基于深度学习的超分辨率技术在资源受限的边缘设备上实现超过全高清分辨率的应用面临高计算复杂度与内存带宽需求的挑战。本文提出一款支持8K@30FPS的超分辨率加速器,采用边缘感知的动态输入处理机制。该机制依据输入图像的边缘特性,动态选择合适的子网络,实现50%的乘加操作(MAC)减少,仅导致0.1dB的PSNR下降。通过资源自适应模型切换,在资源受限条件下仍可保障重建图像质量。结合硬件特化优化,模型尺寸缩减84%至51K,PSNR损失小于0.6dB。为支持高效动态处理,设计了可配置层映射组,与结构友好融合模块协同工作,使硬件利用率提升至77%,特征存储访问减少79%。基于TSMC 28nm工艺实现,运行频率800MHz,门电路数2749K,功耗0.2075W,能效4797Mpixels/J,优于现有工作。

原文摘要 · Abstract (English)

Deep learning-based super-resolution (SR) is challenging to implement in resource-constrained edge devices for resolutions beyond full HD due to its high computational complexity and memory bandwidth requirements. This paper introduces an 8K@30FPS SR accelerator with edge-selective dynamic input processing. Dynamic processing chooses the appropriate subnets for different patches based on simple input edge criteria, achieving a 50\% MAC reduction with only a 0.1dB PSNR decrease. The quality of reconstruction images is guaranteed and maximized its potential with \textit{resource adaptive model switching} even under resource constraints. In conjunction with hardware-specific refinements, the model size is reduced by 84\% to 51K, but with a decrease of less than 0.6dB PSNR. Additionally, to support dynamic processing with high utilization, this design incorporates a \textit{configurable group of layer mapping} that synergizes with the \textit{structure-friendly fusion block}, resulting in 77\% hardware utilization and up to 79\% reduction in feature SRAM access. The implementation, using the TSMC 28nm process, can achieve 8K@30FPS throughput at 800MHz with a gate count of 2749K, 0.2075W power consumption, and 4797Mpixels/J energy efficiency, exceeding previous work.

超分辨率边缘计算硬件加速8K视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。