arXiv:2504.14416cs.LG2025-04被引 2

用伪标记提升Transformer神经过程的效率与可扩展性

Exploring Pseudo-Token Approaches in Transformer Neural Processes

  • 引入诱导集注意力机制,通过伪标记压缩上下文信息
  • 在1D回归等任务中性能接近甚至超越现有最优模型
  • 可调节计算开销,适合大规模数据场景

神经过程(NPs)因其能量化不确定性、快速预测和良好适应性,在元学习中受到关注。但传统NPs易出现欠拟合问题。基于Transformer的神经过程(TNPs)显著优于现有NPs,但其在真实场景中的应用受限于与上下文和目标数据点数量呈二次方增长的计算复杂度。为此,伪标记型TNPs(PT-TNPs)作为新分支被提出,通过将上下文数据压缩为潜在向量或伪标记以降低计算负担。本文提出诱导集注意神经过程(ISANPs),采用诱导集注意力机制与创新查询阶段,提升查询效率。实验表明,ISANPs在1D回归、图像补全、上下文老虎机及贝叶斯优化任务中表现优异,性能与TNPs相当甚至更优。关键优势在于可调节性能与计算复杂度的平衡,使其在大规模数据下仍具可扩展性。

原文摘要 · Abstract (English)

Neural Processes (NPs) have gained attention in meta-learning for their ability to quantify uncertainty, together with their rapid prediction and adaptability. However, traditional NPs are prone to underfitting. Transformer Neural Processes (TNPs) significantly outperform existing NPs, yet their applicability in real-world scenarios is hindered by their quadratic computational complexity relative to both context and target data points. To address this, pseudo-token-based TNPs (PT-TNPs) have emerged as a novel NPs subset that condense context data into latent vectors or pseudo-tokens, reducing computational demands. We introduce the Induced Set Attentive Neural Processes (ISANPs), employing Induced Set Attention and an innovative query phase to improve querying efficiency. Our evaluations show that ISANPs perform competitively with TNPs and often surpass state-of-the-art models in 1D regression, image completion, contextual bandits, and Bayesian optimization. Crucially, ISANPs offer a tunable balance between performance and computational complexity, which scale well to larger datasets where TNPs face limitations.

神经过程注意力机制高效建模元学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。