arXiv:2506.08785cs.ARcs.AI2025-06被引 11

面向边缘设备的可自适应精度计算架构,兼顾能效与模型精度。

POLARON: Precision-aware On-device Learning and Adaptive Runtime-cONfigurable AI acceleration

  • 采用统一数据通路支持多种精度格式,实现多精度乘累加运算。
  • 相比现有设计,能效提升2倍,资源占用减少3倍,精度损失小于1.8%。
  • 支持边缘端训练与推理,适合复杂AI模型部署场景。

随着AI模型复杂度提升,边缘平台亟需支持多样精度格式的灵活硬件。本文提出PARV-CE,一种基于SIMD的多精度乘累加引擎,统一支持4/8/16位定点、浮点及posit格式的高效乘累加操作。通过分层自适应精度策略,使计算精度匹配任务敏感性,优化性能与能耗。PARV-CE融合量化感知执行与可重构SIMD流水线,借助软硬件协同设计实现高吞吐、低开销处理。实验表明,相较最先进设计,该架构在功耗延迟积(PDP)上提升最高达2倍,资源使用减少3倍,精度仅比FP32基线下降1.8%以内。该架构支持在边缘设备上进行包括深度神经网络、循环神经网络、强化学习及Transformer在内的多种模型的训练与推理。实证分析验证了集成POLARON的PARV-CE在边缘端具备可扩展、高能效的精度自适应AI加速能力。

原文摘要 · Abstract (English)

The increasing complexity of AI models requires flexible hardware capable of supporting diverse precision formats, particularly for energy-constrained edge platforms. This work presents PARV-CE, a SIMD-enabled, multi-precision MAC engine that performs efficient multiply-accumulate operations using a unified data-path for 4/8/16-bit fixed-point, floating point, and posit formats. The architecture incorporates a layer adaptive precision strategy to align computational accuracy with workload sensitivity, optimizing both performance and energy usage. PARV-CE integrates quantization-aware execution with a reconfigurable SIMD pipeline, enabling high-throughput processing with minimal overhead through hardware-software co-design. The results demonstrate up to 2x improvement in PDP and 3x reduction in resource usage compared to SoTA designs, while retaining accuracy within 1.8% FP32 baseline. The architecture supports both on-device training and inference across a range of workloads, including DNNs, RNNs, RL, and Transformer models. The empirical analysis establish PARVCE incorporated POLARON as a scalable and energy-efficient solution for precision-adaptive AI acceleration at edge.

边缘计算多精度计算AI加速硬件协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。