arXiv:2502.06007stat.MLcs.LG2025-02被引 3

Transformer在聚类任务中表现如EM算法,理论证明其逼近最优性能。

Transformers versus the EM Algorithm in Multi-class Clustering

  • 将Softmax注意力层与EM算法流程关联,建立理论桥梁。
  • 在足够预训练样本下,达到该问题的最小最大最优率。
  • 适合关注大模型推理能力与无监督学习机制的研究者。

大型语言模型在复杂机器学习任务中展现出强大的推理能力,其核心为Transformer架构。针对此类模型在无监督学习问题上理解不足的现状,本文研究了Transformer在高斯混合模型多分类聚类中的学习保证。通过揭示Softmax注意力层与高斯混合聚类中EM算法流程之间的紧密联系,建立了期望和最大化步骤的近似界,证明了Softmax函数对多元映射的通用逼近能力。除近似保证外,还表明在足够多的预训练样本及特定初始化条件下,Transformer可实现所考虑问题的最小最大最优率。大量模拟实验验证了理论结论,揭示Transformer即使在理论假设之外也表现出强大学习能力,为大模型的强推理能力提供了新视角。

原文摘要 · Abstract (English)

LLMs demonstrate significant inference capacities in complicated machine learning tasks, using the Transformer model as its backbone. Motivated by the limited understanding of such models on the unsupervised learning problems, we study the learning guarantees of Transformers in performing multi-class clustering of the Gaussian Mixture Models. We develop a theory drawing strong connections between the Softmax Attention layers and the workflow of the EM algorithm on clustering the mixture of Gaussians. Our theory provides approximation bounds for the Expectation and Maximization steps by proving the universal approximation abilities of multivariate mappings by Softmax functions. In addition to the approximation guarantees, we also show that with a sufficient number of pre-training samples and an initialization, Transformers can achieve the minimax optimal rate for the problem considered. Our extensive simulations empirically verified our theory by revealing the strong learning capacities of Transformers even beyond the assumptions in the theory, shedding light on the powerful inference capacities of LLMs.

Transformer聚类理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。