arXiv:2502.10463cs.LGcs.AI2025-02ICLR被引 2

用状态空间模型重新理解深层网络层间动态,提升视觉任务表现。

From Layers to States: A State Space Model Perspective to Deep Neural Network Layer Dynamics

  • 将网络层输出视为连续状态,用状态空间模型建模层间信息流动。
  • 在图像分类与检测任务上显著优于基线模型,性能提升明显。
  • 适合关注深层网络机制、长序列建模的视觉算法研究者。

神经网络的深度是其能力的关键因素,更深的模型通常表现更优。为提升表征能力,已有工作致力于层间信息聚合——复用前层信息以增强当前层特征提取。然而,现有方法多基于离散状态视角,在层数增多时难以适用。本文首次将层输出视为连续过程中的状态,引入状态空间模型(SSM)设计深层网络的层聚合机制。受长序列建模进展启发,采用选择性状态空间模型(S6)构建新模块S6LA,可无缝集成于传统CNN或Transformer架构中,形成顺序框架,增强主流视觉网络的表征能力。大量实验表明,S6LA在图像分类与目标检测任务中均取得显著提升,验证了将SSM与现代深度学习技术融合的潜力。

原文摘要 · Abstract (English)

The depth of neural networks is a critical factor for their capability, with deeper models often demonstrating superior performance. Motivated by this, significant efforts have been made to enhance layer aggregation - reusing information from previous layers to better extract features at the current layer, to improve the representational power of deep neural networks. However, previous works have primarily addressed this problem from a discrete-state perspective which is not suitable as the number of network layers grows. This paper novelly treats the outputs from layers as states of a continuous process and considers leveraging the state space model (SSM) to design the aggregation of layers in very deep neural networks. Moreover, inspired by its advancements in modeling long sequences, the Selective State Space Models (S6) is employed to design a new module called Selective State Space Model Layer Aggregation (S6LA). This module aims to combine traditional CNN or transformer architectures within a sequential framework, enhancing the representational capabilities of state-of-the-art vision networks. Extensive experiments show that S6LA delivers substantial improvements in both image classification and detection tasks, highlighting the potential of integrating SSMs with contemporary deep learning techniques.

状态空间模型深层网络视觉任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。