arXiv:2602.19268cs.ARcs.AI2026-02被引 1

用迭代式CORDIC单元实现低资源向量计算,动态调节精度以提升边缘AI性能。

CORVET: A CORDIC-Powered, Resource-Frugal Mixed-Precision Vector Processing Engine for High-Throughput AIoT applications

  • 基于CORDIC的轻量级乘加单元支持运行时精度切换。
  • 相同硬件下吞吐量最高提升4倍,能效达11.67 TOPS/W。
  • 适合对功耗敏感的嵌入式视觉任务,如物体检测与分类。

本文提出一种面向边缘AI加速的可运行时自适应、高性能向量引擎,采用低资源、迭代式CORDIC-based乘加(MAC)单元。该设计支持近似与精确模式之间的动态重构,利用延迟-精度权衡应对多种工作负载。其资源高效方法通过向量化时间复用执行和灵活精度缩放,可在相同硬件资源下实现高达4倍的吞吐量提升。结合时间复用的多-仿射变换块与轻量级池化归一化单元,该向量引擎支持4/8/16位灵活精度和高乘加密度。ASIC实现结果显示,每个MAC阶段最多节省33%时间与21%功耗;256个处理单元配置下,计算密度达4.83 TOPS/mm²,能效达11.67 TOPS/W,优于现有最先进方案。针对Pynq-Z2平台,详细阐述了软硬件协同设计方法,用于物体检测与分类任务,验证了其在边缘AI应用中的可扩展性与能效优势。

原文摘要 · Abstract (English)

This brief presents a runtime-adaptive, performance-enhanced vector engine featuring a low-resource, iterative CORDIC-based MAC unit for edge AI acceleration. The proposed design enables dynamic reconfiguration between approximate and accurate modes, exploiting the latency-accuracy trade-off for a wide range of workloads. Its resource-efficient approach further enables up to 4x throughput improvement within the same hardware resources by leveraging vectorised, time-multiplexed execution and flexible precision scaling. With a time-multiplexed multi-AF block and a lightweight pooling and normalisation unit, the proposed vector engine supports flexible precision (4/8/16-bit) and high MAC density. The ASIC implementation results show that each MAC stage can save up to 33% of time and 21% of power, with a 256-PE configuration that achieves higher compute density (4.83 TOPS/mm2 ) and energy efficiency (11.67 TOPS/W) than previous state-of-the-art work. A detailed hardware-software co-design methodology for object detection and classification tasks on Pynq-Z2 is discussed to assess the proposed architecture, demonstrating a scalable, energy-efficient solution for edge AI applications.

边缘AICORDIC低功耗向量引擎

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。