arXiv:2510.24951cs.LGcs.AR2025-10被引 1

通过软硬件协同设计,实现嵌入式场景下高效可靠的神经网络推理。

Resource-Efficient and Robust Inference of Deep and Bayesian Neural Networks on Embedded and Analog Computing Platforms

  • 基于敏感性分析与硬件反馈的分层自动压缩方法
  • 在模拟硬件上提升鲁棒性,支持非平稳条件下的噪声训练
  • 首次将概率光子计算用于硬件级快速推理,节能且稳定

尽管现代机器学习已广泛应用于多个领域,其日益增长的计算需求正严重制约嵌入式和资源受限平台的可扩展性与效率。实际应用中,神经网络不仅需高效运行,还需在分布偏移或未知数据下提供可靠预测。贝叶斯神经网络虽能量化不确定性,但其计算开销进一步加剧了挑战。本文通过算法与硬件协同优化,推动深度与贝叶斯神经网络的资源高效、鲁棒推理。提出Galen方法,实现基于敏感性分析与硬件反馈的分层自动压缩;针对模拟硬件的噪声问题,建模器件缺陷并扩展噪声训练至非平稳条件,提升稳定性;发展解析与集成近似方法,替代昂贵采样,融入编译栈,优化嵌入式推理;最后探索概率光子计算,利用可控模拟噪声作为内在熵源,实现硬件级快速、低功耗的概率推理。上述工作表明,通过软硬件协同设计可同步提升效率与可靠性,为下一代可信、低功耗机器学习系统奠定基础。

原文摘要 · Abstract (English)

While modern machine learning has transformed numerous application domains, its growing computational demands increasingly constrain scalability and efficiency, particularly on embedded and resource-limited platforms. In practice, neural networks must not only operate efficiently but also provide reliable predictions under distributional shifts or unseen data. Bayesian neural networks offer a principled framework for quantifying uncertainty, yet their computational overhead further compounds these challenges. This work advances resource-efficient and robust inference for both conventional and Bayesian neural networks through the joint pursuit of algorithmic and hardware efficiency. The former reduces computation through model compression and approximate Bayesian inference, while the latter optimizes deployment on digital accelerators and explores analog hardware, bridging algorithmic design and physical realization. The first contribution, Galen, performs automatic layer-specific compression guided by sensitivity analysis and hardware-in-the-loop feedback. Analog accelerators offer efficiency gains at the cost of noise; this work models device imperfections and extends noisy training to nonstationary conditions, improving robustness and stability. A second line of work advances probabilistic inference, developing analytic and ensemble approximations that replace costly sampling, integrate into a compiler stack, and optimize embedded inference. Finally, probabilistic photonic computing introduces a paradigm where controlled analog noise acts as an intrinsic entropy source, enabling fast, energy-efficient probabilistic inference directly in hardware. Together, these studies demonstrate how efficiency and reliability can be advanced jointly through algorithm-hardware co-design, laying the foundation for the next generation of trustworthy, energy-efficient machine-learning systems.

神经网络边缘计算概率推理光子计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。