综述嵌入式系统中高效深度学习架构,助力智能设备落地。
Efficient Deep Learning Infrastructures for Embedded Computing Systems: A Comprehensive Survey and Future Envision
- 从手动设计到自动优化,覆盖模型压缩与软硬件协同
- 涵盖从传统CNN到大语言模型的全链条高效方案
- 适合研究嵌入式AI、边缘计算及资源受限场景的开发者
深度神经网络(DNN)在图像分类、目标检测、跟踪与分割等众多视觉和语言处理任务中取得了显著成功。然而,现有主流DNN不断变深变宽,导致训练与推理所需的计算资源急剧增加,进一步拉大了计算密集型DNN与资源受限嵌入式系统之间的差距,使得在真实嵌入式设备上部署高性能DNN面临挑战。为缓解这一计算鸿沟并推动无处不在的嵌入式智能,本文全面综述面向嵌入式计算系统的高效深度学习基础设施,涵盖从训练到推理、从人工设计到自动化设计、从卷积神经网络到Transformer、从视觉模型到大语言模型、从软件到硬件、从算法到应用的全栈技术。具体包括:(1)面向嵌入式系统的高效人工网络设计;(2)高效自动网络设计;(3)网络压缩技术;(4)设备端学习;(5)嵌入式大语言模型;(6)软硬件协同优化;(7)智能应用部署。旨在为实现嵌入式环境下的高效智能提供系统性指引。
原文摘要 · Abstract (English)
Deep neural networks (DNNs) have recently achieved impressive success across a wide range of real-world vision and language processing tasks, spanning from image classification to many other downstream vision tasks, such as object detection, tracking, and segmentation. However, previous well-established DNNs, despite being able to maintain superior accuracy, have also been evolving to be deeper and wider and thus inevitably necessitate prohibitive computational resources for both training and inference. This trend further enlarges the computational gap between computation-intensive DNNs and resource-constrained embedded computing systems, making it challenging to deploy powerful DNNs upon real-world embedded computing systems towards ubiquitous embedded intelligence. To alleviate the above computational gap and enable ubiquitous embedded intelligence, we, in this survey, focus on discussing recent efficient deep learning infrastructures for embedded computing systems, spanning from training to inference, from manual to automated, from convolutional neural networks to transformers, from transformers to vision transformers, from vision models to large language models, from software to hardware, and from algorithms to applications. Specifically, we discuss recent efficient deep learning infrastructures for embedded computing systems from the lens of (1) efficient manual network design for embedded computing systems, (2) efficient automated network design for embedded computing systems, (3) efficient network compression for embedded computing systems, (4) efficient on-device learning for embedded computing systems, (5) efficient large language models for embedded computing systems, (6) efficient deep learning software and hardware for embedded computing systems, and (7) efficient intelligent applications for embedded computing systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。