arXiv:2505.22158cs.LG2025-05

揭示梯度信息量不足的根源,给出可量化评估的理论边界。

The informativeness of the gradient revisited

  • 从函数类对称性与输入分布碰撞熵出发,推导梯度方差上界。
  • 理论表明梯度信息随输入分布碰撞熵指数衰减,影响优化效率。
  • 适用于分析密码学攻击中的深度学习方法,指导模型设计。

过去十年,基于梯度的深度学习在多个应用中取得突破性进展。然而,这一快速发展也凸显出对其局限性的深层理论理解的迫切需求。研究表明,在许多实际学习任务中,梯度所包含的信息极为有限,导致梯度方法需要极多迭代才能成功。梯度的资讯量通常通过目标函数在假设类中随机选取时的方差来衡量。本文采用该框架,给出了一个关于目标函数类成对独立性参数和输入分布碰撞熵的通用方差上界,其形式为 $ \tilde{\mathcal{O}}(\varepsilon + e^{-\frac{1}{2}\mathcal{E}_c}) $,其中 $ \tilde{\mathcal{O}} $ 隐含了学习模型和损失函数的正则性因子,$ \varepsilon $ 衡量目标函数类的成对独立性,$ \mathcal{E}_c $ 为输入分布的碰撞熵。为验证该边界的实用性,我们将其应用于学习误差(LWE)映射和高频函数。此外,还通过实验进一步理解近期基于深度学习的LWE攻击机制。

原文摘要 · Abstract (English)

In the past decade gradient-based deep learning has revolutionized several applications. However, this rapid advancement has highlighted the need for a deeper theoretical understanding of its limitations. Research has shown that, in many practical learning tasks, the information contained in the gradient is so minimal that gradient-based methods require an exceedingly large number of iterations to achieve success. The informativeness of the gradient is typically measured by its variance with respect to the random selection of a target function from a hypothesis class. We use this framework and give a general bound on the variance in terms of a parameter related to the pairwise independence of the target function class and the collision entropy of the input distribution. Our bound scales as $ \tilde{\mathcal{O}}(\varepsilon+e^{-\frac{1}{2}\mathcal{E}_c}) $, where $ \tilde{\mathcal{O}} $ hides factors related to the regularity of the learning model and the loss function, $ \varepsilon $ measures the pairwise independence of the target function class and $\mathcal{E}_c$ is the collision entropy of the input distribution. To demonstrate the practical utility of our bound, we apply it to the class of Learning with Errors (LWE) mappings and high-frequency functions. In addition to the theoretical analysis, we present experiments to understand better the nature of recent deep learning-based attacks on LWE.

梯度分析理论深度学习密码学攻击碰撞熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。