arXiv:2512.04051cs.LGmath.OC2025-12

提出离散更新规则,实现高效低精度训练

Convergence for Discrete Parameter Update Schemes

  • 更新规则本身为离散形式,避免连续更新的量化过程
  • 证明了通用离散更新方案的收敛性,具理论保障
  • 适合结构天然离散的模型,如神经符号系统

现代深度学习模型需要大量计算资源,推动了低精度训练的研究。量化训练通过将训练组件表示为低比特整数来缓解这一问题,但通常依赖对实值更新进行离散化。本文提出一种新思路:直接让更新规则本身为离散形式,从设计上避免连续更新的量化。我们建立了此类离散更新方案的通用收敛性保证,并以多项式更新规则为例进行了实证评估。该视角为高效训练开辟了新路径,尤其适用于具有固有离散结构的模型。

原文摘要 · Abstract (English)

Modern deep learning models require immense computational resources, motivating research into low-precision training. Quantised training addresses this by representing training components in low-bit integers, but typically relies on discretising real-valued updates. We introduce an alternative approach where the update rule itself is discrete, avoiding the quantisation of continuous updates by design. We establish convergence guarantees for a general class of such discrete schemes, and present a multinomial update rule as a concrete example, supported by empirical evaluation. This perspective opens new avenues for efficient training, particularly for models with inherently discrete structure.

低精度训练离散更新收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。