用博弈论解释深度学习为何偏爱捷径特征。
Deciphering Shortcut Learning from an Evolutionary Game Theory Perspective

- 将数据样本视为玩家,特征作为策略,构建演化博弈模型。
- 梯度下降易锁定捷径子网络,随机梯度下降则倾向核心特征。
- 揭示噪声如何诱发捷径偏差,为缓解提供理论依据。
捷径学习导致深度学习模型依赖数据中的非关键特征,但其在神经网络训练中的形成机制仍缺乏理论解释。本文首次形式化定义核心特征与捷径特征,并运用演化博弈论建模:将数据样本视为参与者,其对应的神经正切特征作为策略,假设存在核心与捷径子网络。研究发现,梯度下降(GD)和随机梯度下降(SGD)分别导向两种不同的随机稳定状态,前者主要优化捷径子网络,后者则聚焦核心子网络。通过连续随机微分方程分析策略影响,揭示数据噪声与优化噪声对捷径偏差形成的决定性作用。本工作首次以演化博弈论刻画捷径偏差的动态演化过程,为理解并缓解该现象提供了理论框架。
原文摘要 · Abstract (English)
Shortcut learning causes deep learning models to rely on non-essential features within the data. However, its formation in deep neural network training still lacks theoretical understanding. In this paper, we provide a formal definition of core and shortcut features and employ evolutionary game theory to analyze the origins of shortcut bias by modeling data samples as players and their corresponding neural tangent features as strategies, assuming the existence of core and shortcut subnetworks. We find that gradient descent (GD) and stochastic gradient descent (SGD) lead to two distinct stochastically stable states, each corresponding to a different strategy. The former primarily optimizes the shortcut subnetwork, while the latter primarily optimizes the core subnetwork. We investigate the influence of these strategies on shortcut bias through a continuous stochastic differential equation, and reveal the impact of data noise and optimization noise on the formation of shortcut bias. In brief, our work employs evolutionary game theory to characterize the dynamics of shortcut bias formation and provides a theoretical view on its mitigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。