arXiv:2512.09275stat.MLcs.LG2025-12被引 2

首次揭示位置编码会放大Transformer的泛化误差和对抗脆弱性

Impact of Positional Encoding: Clean and Adversarial Rademacher Complexity for Transformers under In-Context Regression

  • 构建可训练位置编码的单层Transformer泛化分析框架
  • 发现位置编码使干净与对抗场景下的泛化差距均扩大
  • 适合关注模型鲁棒性与提示学习泛化机制的研究者

位置编码(PE)是Transformer的核心组件,但其对模型泛化与鲁棒性的影响尚不明确。本文首次针对可训练位置编码模块,在上下文学习回归任务下提供单层Transformer的泛化分析。结果表明,位置编码系统性地增大了泛化误差。进一步扩展至对抗设置,推导出对抗Rademacher泛化界,发现有无位置编码的模型在攻击下差距被放大,证明位置编码加剧了模型的脆弱性。实验模拟验证了理论边界。本工作建立了一个理解含位置编码的上下文学习中干净与对抗泛化的全新框架。

原文摘要 · Abstract (English)

Positional encoding (PE) is a core architectural component of Transformers, yet its impact on the Transformer's generalization and robustness remains unclear. In this work, we provide the first generalization analysis for a single-layer Transformer under in-context regression that explicitly accounts for a completely trainable PE module. Our result shows that PE systematically enlarges the generalization gap. Extending to the adversarial setting, we derive the adversarial Rademacher generalization bound. We find that the gap between models with and without PE is magnified under attack, demonstrating that PE amplifies the vulnerability of models. Our bounds are empirically validated by a simulation study. Together, this work establishes a new framework for understanding the clean and adversarial generalization in ICL with PE.

Transformer位置编码泛化分析对抗鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。