arXiv:2502.03327stat.MLcs.LG2025-02被引 4

MLP也能实现上下文通用,未必是模型结构优势

Is In-Context Universality Enough? MLPs are Also Universal In-Context

  • 用可训练激活函数的MLP实现上下文通用性
  • 证明了非Transformer架构同样具备函数逼近能力
  • 适合研究模型泛化能力与结构优势的学者

Transformer的成功常归因于其上下文学习能力。近期研究表明,Transformer具有上下文通用性,能够逼近任意实值连续函数,该函数依赖于上下文(定义在$/mathcal{X}볮q bR^d$上的概率测度)和查询点 $x볮q /mathcal{X}$。这引发疑问:上下文通用性是否解释了其相对于经典模型的优势?我们通过证明:具有可训练激活函数的MLP同样具备上下文通用性,从而否定这一假设。这表明Transformer的成功更可能源于归纳偏置或训练稳定性等其他因素。

原文摘要 · Abstract (English)

The success of transformers is often linked to their ability to perform in-context learning. Recent work shows that transformers are universal in context, capable of approximating any real-valued continuous function of a context (a probability measure over $\mathcal{X}\subseteq \mathbb{R}^d$) and a query $x\in \mathcal{X}$. This raises the question: Does in-context universality explain their advantage over classical models? We answer this in the negative by proving that MLPs with trainable activation functions are also universal in-context. This suggests the transformer's success is likely due to other factors like inductive bias or training stability.

MLP上下文学习通用性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。