提出软硬件协同设计框架,让生成式视觉AI在边缘设备更高效可用。
Vision-centric generative AI models: A software-hardware perspective

- 从模型设计之初就考虑部署约束,实现软硬件协同优化。
- 量化对比多类生成模型在不同加速器上的参数量与能效表现。
- 适合关注边缘计算、移动设备部署的AI工程师与系统设计师。
视觉生成人工智能已成为深度学习中发展最快的领域之一。多模态模型的爆发推动了文本到图像应用在大型数据中心的普及,但生成式视觉模型在自动驾驶、农业传感器和移动设备等硬件受限的边缘场景中同样至关重要。本文指出,当前视觉生成AI的进步主要依赖输出质量提升,而硬件则被动适应模型增长需求。我们量化了多种加速器平台上生成模型的参数开销与能效表现,并将四类生成模型家族映射至七个实际应用场景。最后,我们倡导软件-硬件协同设计方法,在设计初期即纳入部署约束,确保‘合适的模型’运行在‘合适的硬件’上,服务于‘合适的场景’,从而实现生成式AI在更广泛平台上的可持续与可及性部署。
原文摘要 · Abstract (English)
Vision generative artificial intelligence (AI) has emerged as one of the most rapidly advancing areas of deep learning. The explosion of multimodal models has made them widely associated with text-to-image applications running on large datacentres. However, vision generative models are equally needed in applications that operate under strict hardware constraints at the edge, including autonomous vehicles, agricultural sensors, and mobile devices. In this Perspective, we argue that progress in vision generative AI has been driven by output quality, with hardware evolving reactively to accommodate growing model demands. We quantify the parameter cost and energy efficiency of these models across a range of accelerator platforms, and map four generative model families against seven real-world application domains. Finally, we advocate a software-hardware co-design approach, where deployment constraints are considered from the start of the design process, ensuring that the "right model" runs on the "right hardware" to serve the "right application", making generative AI deployment sustainable and accessible across a much broader range of platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。