arXiv:2509.17514cs.LG2025-09NeurIPS被引 4

用合成数据揭示Mamba架构的对称性识别缺陷

Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data

  • 通过合成任务暴露Mamba非线性卷积导致的信息融合不对称
  • 在逆序匹配任务中表现显著落后,无法有效识别对称关系
  • 适合研究序列模型局限性或设计更对称的架构者阅读

状态空间模型(SSMs)已成为注意力机制的有前途替代方案,Mamba架构在处理长序列时展现出优异性能与线性复杂度。然而,Mamba与Transformer架构的根本差异仍不明确。本文通过精心设计的合成任务揭示Mamba的固有局限。实验表明,Mamba中的非线性卷积引入了不对称偏差,严重削弱其识别对称模式与关系的能力。在复合函数和逆序匹配任务中,Mamba强烈偏好组合解而非对称解,且难以完成反向序列匹配。这些限制并非源于SSM模块本身,而是由其前的非线性卷积造成,该模块以不对称方式融合标记信息。这些发现为理解Mamba的约束提供了新视角,并为未来序列模型的改进提供了具体方向。

原文摘要 · Abstract (English)

State Space Models (SSMs) have emerged as promising alternatives to attention mechanisms, with the Mamba architecture demonstrating impressive performance and linear complexity for processing long sequences. However, the fundamental differences between Mamba and Transformer architectures remain incompletely understood. In this work, we use carefully designed synthetic tasks to reveal Mamba's inherent limitations. Through experiments, we identify that Mamba's nonlinear convolution introduces an asymmetry bias that significantly impairs its ability to recognize symmetrical patterns and relationships. Using composite function and inverse sequence matching tasks, we demonstrate that Mamba strongly favors compositional solutions over symmetrical ones and struggles with tasks requiring the matching of reversed sequences. We show these limitations stem not from the SSM module itself but from the nonlinear convolution preceding it, which fuses token information asymmetrically. These insights provide a new understanding of Mamba's constraints and suggest concrete architectural improvements for future sequence models.

序列建模Mamba对称性合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。