用FPGA加速医疗数字孪生学习,显著提升能效与速度。
Accelerated Digital Twin Learning for Edge AI: A Comparison of FPGA and Mobile GPU
- 将数字孪生学习框架适配可重构硬件FPGA,实现高效计算。
- 相比云端GPU,FPGA提升8.8倍能效、减少28.5%内存占用、提速1.67倍。
- 适用于糖尿病和冠心病等边缘AI医疗场景的实时建模与数据生成。
数字孪生(DT)可通过持续学习患者特异性动态的数学表征,实现精准医疗。然而,关键医疗应用需快速、低资源的DT学习,现有模型恢复(MR)技术因依赖迭代求解器且计算/内存需求高,难以满足。本文提出一种适用于可重构硬件(如FPGA)的通用DT学习框架,实现显著加速与能效提升。我们对比了FPGA实现与移动GPU多进程实现,并与云GPU基线进行比较。结果表明,FPGA在MR任务上实现8.8倍性能-功耗比提升,DRAM占用降低28.5%,运行时间快1.67倍;而移动GPU虽性能-功耗比优2倍,但运行时间增加2倍,内存占用高出FPGA 10倍。该方法成功应用于1型糖尿病合成数据生成与冠心病早期预警。
原文摘要 · Abstract (English)
Digital twins (DTs) can enable precision healthcare by continually learning a mathematical representation of patient-specific dynamics. However, mission critical healthcare applications require fast, resource-efficient DT learning, which is often infeasible with existing model recovery (MR) techniques due to their reliance on iterative solvers and high compute/memory demands. In this paper, we present a general DT learning framework that is amenable to acceleration on reconfigurable hardware such as FPGAs, enabling substantial speedup and energy efficiency. We compare our FPGA-based implementation with a multi-processing implementation in mobile GPU, which is a popular choice for AI in edge devices. Further, we compare both edge AI implementations with cloud GPU baseline. Specifically, our FPGA implementation achieves an 8.8x improvement in \text{performance-per-watt} for the MR task, a 28.5x reduction in DRAM footprint, and a 1.67x runtime speedup compared to cloud GPU baselines. On the other hand, mobile GPU achieves 2x better performance per watts but has 2x increase in runtime and 10x more DRAM footprint than FPGA. We show the usage of this technique in DT guided synthetic data generation for Type 1 Diabetes and proactive coronary artery disease detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。