用数学框架解释变压器模型在语义分布外时为何失效
A Theoretical Framework for OOD Robustness in Transformers using Gevrey Classes
- 基于沃瑟斯坦距离与盖弗雷类平滑性,推导出预测误差的次指数上界
- 实验显示算术与思维链任务中,分布偏移导致性能下降,符合理论预测
- 适合研究模型鲁棒性与泛化机制的学者,尤其关注分布外场景
我们研究了在语义分布外(OOD)迁移下,Transformer语言模型的鲁棒性,此时训练与测试数据位于不相交的隐空间中。利用一阶沃瑟斯坦距离与盖弗雷类平滑性,我们推导出预测误差的次指数上界。该理论框架揭示了平滑性如何在分布漂移下决定泛化能力。通过在算术与链式思维任务上进行受控实验,对隐空间排列与缩放进行操作,结果表明实际性能退化与理论边界一致,凸显了变压器模型在分布外泛化中的几何与函数原理。
原文摘要 · Abstract (English)
We study the robustness of Transformer language models under semantic out-of-distribution (OOD) shifts, where training and test data lie in disjoint latent spaces. Using Wasserstein-1 distance and Gevrey-class smoothness, we derive sub-exponential upper bounds on prediction error. Our theoretical framework explains how smoothness governs generalization under distributional drift. We validate these findings through controlled experiments on arithmetic and Chain-of-Thought tasks with latent permutations and scalings. Results show empirical degradation aligns with our bounds, highlighting the geometric and functional principles underlying OOD generalization in Transformers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。