arXiv:2501.01312stat.MLcs.LG2025-01被引 3

Transformer可自学谱方法,实现无监督统计估计。

Learning Spectral Methods by Transformers

  • 用多层Transformer预训练学习谱方法算法
  • 在合成与真实数据上成功完成主成分分析和聚类
  • 适合研究模型自学习机制的学者参考

Transformer在现代大模型中表现出显著优势。本文研究其在无监督学习中的能力,发现多层Transformer在足够大的预训练实例集下,能自主学习算法并完成新实例上的统计估计任务。该范式不同于上下文学习,更类似于人类通过经验习得技能。理论上,我们证明了预训练Transformer可学习谱方法,并以双类高斯混合模型分类为例,采用算法设计技术给出构造性证明。理论基础源于多层Transformer架构与实际迭代恢复算法的相似性。实验上,我们在合成与真实数据集上验证了多层(预训练)Transformer在主成分分析与聚类任务中的强大无监督学习能力。

原文摘要 · Abstract (English)

Transformers demonstrate significant advantages as the building block of modern LLMs. In this work, we study the capacities of Transformers in performing unsupervised learning. We show that multi-layered Transformers, given a sufficiently large set of pre-training instances, are able to learn the algorithms themselves and perform statistical estimation tasks given new instances. This learning paradigm is distinct from the in-context learning setup and is similar to the learning procedure of human brains where skills are learned through past experience. Theoretically, we prove that pre-trained Transformers can learn the spectral methods and use the classification of bi-class Gaussian mixture model as an example. Our proof is constructive using algorithmic design techniques. Our results are built upon the similarities of multi-layered Transformer architecture with the iterative recovery algorithms used in practice. Empirically, we verify the strong capacity of the multi-layered (pre-trained) Transformer on unsupervised learning through the lens of both the PCA and the Clustering tasks performed on the synthetic and real-world datasets.

Transformer无监督学习谱方法自学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。