提出快速非奇异的熵近似方法,显著提升计算效率与模型性能。
Fast, close, non-singular and property-preserving approximations of entropic measures
- 用有理函数近似熵和对称KL散度,避免梯度奇异问题。
- 平均绝对误差约10^-3,计算速度比现有方法快2至37倍。
- 适合需要高效熵计算的机器学习与量子计算场景。
香农熵(SE)、冯诺依曼熵及其对称KL散度是物理、信息论、机器学习和量子计算中许多工具的核心。这些度量在实际应用中常需大量计算,且其在零邻域的梯度奇异导致相关算法计算成本高、鲁棒性差、收敛慢。本文提出快速熵近似(FEA)——一种非奇异的有理近似方法,用于SE和对称KL散度,保留其主要数学性质,平均绝对误差约为10^-3(比现有先进方法优10-20倍)。实验表明,FEA使SE计算提速约2倍,对称KL计算提速高达37倍,仅需5至7个基本运算操作,远少于依赖查表、位移或级数逼近的现有方案。在典型机器学习特征选择基准上,结合低误差、少操作数、保性质与非奇异梯度,训练速度大幅提升,模型质量显著优于传统方法。结果显示,基于FEA的特征提取比广泛使用的LASSO快三个数量级,且质量更优。
原文摘要 · Abstract (English)
Entropic measures like Shannon entropy (SE), its quantum mechanical analogue von Neumann entropy, and Kullback-Leibler divergence (KL) are key components in many tools used in physics, information theory, machine learning (ML) and quantum computing. Besides of the significant amounts of SE and KL computations required in these fields, the singularity of their gradients near zero is one of the central mathematical reason inducing the high cost, frequently low robustness and slow convergence of computational tools that rely on these concepts. Here we propose the Fast Entropic Approximations (FEA) - non-singular rational approximations of SE and symmetrized KL, that preserve their main mathematical properties and achieve a mean absolute errors of around $10^-3$ ($10-20$ times better than comparable state-of-the-art computational approximations). We show that FEA allows up to around 2 times faster computation of SE and up to 37 times faster computation of symmetrized KL: it requires only $5$ to $7$ elementary computational operations, as compared to the tens of elementary operations behind SE and KL evaluations based on approximate logarithm schemes with table look-ups, bitshifts, or series approximations. On a set of common benchmarks for the feature selection problem in machine learning, we show that the combined effect of fewer elementary operations, low approximation error, preservation of main mathematical properties, and non-singular gradients allows much faster training of significantly-better models. We demonstrate that FEA enables ML feature extraction that is three orders of magnitude faster, and better in quality then the very popular LASSO feature extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。