分析六类硬件畸变对训练补偿能力的影响,区分可调优与不可调优的误差类型。
A Controlled Diagnostic Study of Hardware-Induced Distortions in Hardware-Aware Training
- 将硬件非理想性建模为前向算子的结构化扰动,用梯度优化兼容性评估其可补偿性。
- 发现读噪声、变异、漂移等三类畸变可被训练补偿,而滞留故障等三类会破坏优化过程。
- 为软硬件协同设计提供依据:明确哪些误差需在电路或校准层面解决。
硬件感知训练(HAT)广泛用于提升神经网络在非理想AI加速器(如模拟存内计算系统)上的鲁棒性。然而,并非所有硬件引起的畸变都能通过训练有效补偿。本文提出一种诊断框架,将硬件非理想性建模为前向算子的结构化扰动,评估其与基于梯度优化的兼容性。分析六类典型扰动——读噪声、变异、漂移、滞留故障、IR压降和ADC量化——并识别出三个关键诊断指标:梯度期望一致性、有界梯度方差和非退化敏感性。结果表明,可被HAT补偿的扰动与会持续破坏优化的扰动之间存在明显分界。该研究为软硬件协同设计提供了实践指导,明确了哪些非理想性可在训练阶段处理,哪些必须通过电路、架构或校准手段缓解。本研究为标准前向扰动型HAT下的受控实证分析,而非通用硬件感知训练理论。
原文摘要 · Abstract (English)
Hardware-aware training (HAT) is widely used to improve the robustness of neural networks on non-ideal AI accelerators, such as analog in-memory computing (IMC) systems. However, not all hardware-induced distortions are equally compensable by training. This paper presents a diagnostic framework that models hardware non-idealities as structured perturbations of the forward operator and evaluates their compatibility with gradient-based optimization. We analyze six representative perturbation classes--read noise, variability, drift, stuck-at faults, IR-drop, and ADC discretization--and identify three key diagnostics: gradient expectation consistency, bounded gradient variance, and non-degenerate sensitivity. Our results show a clear separation between perturbations that can be compensated by HAT and those that consistently break optimization. This provides practical guidance for hardware-software co-design, clarifying which non-idealities can be addressed at the training level and which require circuit-, architecture-, or calibration-level mitigation. This study should be interpreted as a controlled empirical analysis under vanilla forward-perturbation HAT, rather than as a universal theory of hardware-aware training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。