揭示Transformer如何在噪声数据中利用低维任务结构实现高效学习
Transformers for Learning on Noisy and Task-Level Manifolds: Approximation and Generalization Insights
- 将输入数据建模为流形邻域内的噪声点,任务依赖于其投影到任务流形的结构
- 理论证明了模型误差与任务流形内在维度密切相关,且在高维噪声下仍能保持性能
- 提出新证明方法,通过Transformer构建基础算术运算表示,具有独立研究价值
Transformer是GPT、BERT、SORA等大语言与视频生成模型的基础架构。实证研究表明,真实世界数据与学习任务具有低维结构,同时存在噪声或测量误差。尽管变压器性能受数据/任务内在维度影响显著,但针对其理论机制的研究仍不充分。本文首次建立理论框架,分析变压器在涉及噪声输入数据的回归任务中的表现。输入数据位于流形的管状邻域内,而真实函数依赖于噪声数据在任务流形上的投影。理论推导出逼近误差与泛化误差,其关键取决于任务流形的内在维度。结果表明,即使输入数据受高维噪声干扰,变压器仍能有效利用任务的低复杂度结构。本文提出的新型证明技术通过变压器构建基本算术运算的表示,可能具有独立研究意义。
原文摘要 · Abstract (English)
Transformers serve as the foundational architecture for large language and video generation models, such as GPT, BERT, SORA and their successors. Empirical studies have demonstrated that real-world data and learning tasks exhibit low-dimensional structures, along with some noise or measurement error. The performance of transformers tends to depend on the intrinsic dimension of the data/tasks, though theoretical understandings remain largely unexplored for transformers. This work establishes a theoretical foundation by analyzing the performance of transformers for regression tasks involving noisy input data near a manifold. Specifically, the input data are in a tubular neighborhood of a manifold, while the ground truth function depends on the projection of the noisy data onto this manifold, referred to as the task-level manifold. We prove approximation and generalization errors which crucially depend on the intrinsic dimension of the task-level manifold. Our results demonstrate that transformers can leverage low-complexity structures in learning task even when the input data are perturbed by high-dimensional noise. Our novel proof technique constructs representations of basic arithmetic operations by transformers, which may hold independent interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。