arXiv:2603.23668cs.ARcs.LG2026-03中稿 · as a poster presen…被引 1

从微型设备到大模型,如何让机器学习更省电。

Energy Efficient Software Hardware CoDesign for Machine Learning: From TinyML to Large Language Models

  • 软硬件协同设计,优化数据流动与内存使用。
  • 覆盖从边缘推理到数据中心的全场景能效方案。
  • 适合关注绿色AI、系统优化的研究者和工程师。

从毫瓦级的TinyML设备到大型语言模型,机器学习在多平台的快速部署使得能效成为可持续AI的主要约束。在不同规模下,性能和能效日益受限于数据移动与内存系统行为,而不仅仅是算术吞吐量。本文综述了涵盖边缘推理与训练到数据中心级大模型服务的能效软硬件协同设计方法,包括加速器架构(如ASIC/FPGA数据流、存算一体设计)与系统级技术(如分块、量化、调度与运行时自适应)。我们提炼出通用的设计权衡与关键因素,并指出普遍存在的短板:跨平台泛化能力弱、协同设计搜索空间大且成本高、工作负载与部署环境间基准不一致。最后,提出分层分解视角,将优化策略映射到计算角色,支持渐进式适应,为构建能耗与碳排放感知的机器学习系统提供实践指导。

原文摘要 · Abstract (English)

The rapid deployment of machine learning across platforms from milliwatt-class TinyML devices to large language models has made energy efficiency a primary constraint for sustainable AI. Across these scales, performance and energy are increasingly limited by data movement and memory-system behavior rather than by arithmetic throughput alone. This work reviews energy efficient software hardware codesign methods spanning edge inference and training to datacenter-scale LLM serving, covering accelerator architectures (e.g., ASIC/FPGA dataflows, processing-/compute-in-memory designs) and system-level techniques (e.g., partitioning, quantization, scheduling, and runtime adaptation). We distill common design levers and trade-offs, and highlight recurring gaps including limited cross-platform generalization, large and costly co-design search spaces, and inconsistent benchmarking across workloads and deployment settings. Finally, we outline a hierarchical decomposition perspective that maps optimization strategies to computational roles and supports incremental adaptation, offering practical guidance for building energy and carbon aware ML systems.

能效优化软硬件协同绿色AI边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。