arXiv:2510.16356cs.LGmath.OC2025-10被引 1

用最优传输设计稀疏Transformer,提升生成模型精度与收敛速度。

Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with $L_1$ Prior

  • 基于正则化Wasserstein近端算子设计稀疏结构
  • 生成样本更稀疏,优化问题凸性更强,收敛更快
  • 适合需要高效生成和逆问题求解的场景

本文提出一种将数据分布先验信息直接融入神经网络Transformer结构的稀疏Transformer架构。该设计源于特定的最优传输问题——正则化Wasserstein近端算子,其具有闭式解,并可视为Transformer的一种特殊形式。相较于经典基于流的模型,该方法提升了优化问题的凸性并促进生成样本的稀疏性。通过理论分析与数值实验(包括生成建模与贝叶斯逆问题应用),结果表明,该稀疏Transformer在逼近目标分布时具有更高精度且收敛速度更快,优于传统的神经ODE方法。

原文摘要 · Abstract (English)

In this work, we propose a sparse transformer architecture that incorporates prior information about the underlying data distribution directly into the transformer structure of the neural network. The design of the model is motivated by a special optimal transport problem, namely the regularized Wasserstein proximal operator, which admits a closed-form solution and turns out to be a special representation of transformer architectures. Compared with classical flow-based models, the proposed approach improves the convexity properties of the optimization problem and promotes sparsity in the generated samples. Through both theoretical analysis and numerical experiments, including applications in generative modeling and Bayesian inverse problems, we demonstrate that the sparse transformer achieves higher accuracy and faster convergence to the target distribution than classical neural ODE-based methods.

Transformer生成模型最优传输稀疏性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。