arXiv:2505.05409cs.LG2025-05ICML被引 6

重新定义变压器模型的尖锐度,揭示其泛化能力的真实规律

Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds It

  • 基于黎曼几何在对称性商流形上重新定义尖锐度
  • 高阶近似下尖锐度与泛化性能相关性显著提升
  • 适用于文本与图像分类任务的Transformer模型分析

尖锐度概念在传统网络如MLP和CNN中成功用于预测泛化性能。然而,对于Transformer,近期研究发现平坦度与泛化之间的相关性较弱。我们认为现有尖锐度度量对Transformer失效,因其注意力机制具有更丰富的对称性,导致参数空间中存在使网络或损失保持不变的方向。因此,我们提出在去除变压器对称性的商流形上重新定义尖锐度,消除歧义。利用黎曼几何工具,我们以对称性修正后的商流形上的测地球来形式化尖锐度。实践中需近似测地线:一阶近似对应现有自适应尖锐度度量,而包含高阶项才可恢复与泛化的强相关性。我们在带合成数据的对角网络上验证方法,并展示在真实世界Transformer(文本与图像分类)上的显著效果。

原文摘要 · Abstract (English)

The concept of sharpness has been successfully applied to traditional architectures like MLPs and CNNs to predict their generalization. For transformers, however, recent work reported weak correlation between flatness and generalization. We argue that existing sharpness measures fail for transformers, because they have much richer symmetries in their attention mechanism that induce directions in parameter space along which the network or its loss remain identical. We posit that sharpness must account fully for these symmetries, and thus we redefine it on a quotient manifold that results from quotienting out the transformer symmetries, thereby removing their ambiguities. Leveraging tools from Riemannian geometry, we propose a fully general notion of sharpness, in terms of a geodesic ball on the symmetry-corrected quotient manifold. In practice, we need to resort to approximating the geodesics. Doing so up to first order yields existing adaptive sharpness measures, and we demonstrate that including higher-order terms is crucial to recover correlation with generalization. We present results on diagonal networks with synthetic data, and show that our geodesic sharpness reveals strong correlation for real-world transformers on both text and image classification tasks.

Transformer尖锐度黎曼几何泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。