arXiv:2603.13663cs.LG2026-03

用物理方程替代注意力,让图像生成模型更高效

PDE-SSM: A Spectral State Space Approach to Spatial Mixing in Diffusion Transformers

  • 用可学习的偏微分方程建模信息流动,替代传统注意力
  • 在傅里叶域求解,实现近线性复杂度 $O(N \log N)$
  • 适合追求高效、强空间先验的视觉生成模型研究者

视觉变换器在生成建模中的成功受限于自注意力机制的二次计算开销和弱空间归纳偏置。我们提出PDE-SSM,一种基于可学习对流-扩散-反应偏微分方程的空间状态空间模块,替代自注意力。该算子通过物理驱动的动力学建模信息流动,编码强空间先验。在傅里叶域求解该方程可实现全局耦合,复杂度接近线性 $O(N \log N)$,提供了一种有原则且可扩展的注意力替代方案。我们将PDE-SSM集成到流匹配生成模型中,构建了基于PDE的扩散变换器PDE-SSM-DiT。实验表明,PDE-SSM-DiT在性能上达到或超越当前最先进扩散变换器,同时显著降低计算开销。结果表明,如同一维场景中状态空间模型取代注意力,多维偏微分算子为下一代视觉模型提供了高效且富含归纳偏置的基础。

原文摘要 · Abstract (English)

The success of vision transformers-especially for generative modeling-is limited by the quadratic cost and weak spatial inductive bias of self-attention. We propose PDE-SSM, a spatial state-space block that replaces attention with a learnable convection-diffusion-reaction partial differential equation. This operator encodes a strong spatial prior by modeling information flow via physically grounded dynamics rather than all-to-all token interactions. Solving the PDE in the Fourier domain yields global coupling with near-linear complexity of $O(N \log N)$, delivering a principled and scalable alternative to attention. We integrate PDE-SSM into a flow-matching generative model to obtain the PDE-based Diffusion Transformer PDE-SSM-DiT. Empirically, PDE-SSM-DiT matches or exceeds the performance of state-of-the-art Diffusion Transformers while substantially reducing compute. Our results show that, analogous to 1D settings where SSMs supplant attention, multi-dimensional PDE operators provide an efficient, inductive-bias-rich foundation for next-generation vision models.

扩散模型状态空间图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。