arXiv:2507.15798cs.CV2025-07被引 1

研究小模型中特征干扰问题,提升低参数视觉网络的准确率与可扩展性。

Exploring Superposition and Interference in State-of-the-Art Low-Parameter Vision Models

  • 通过分析瓶颈结构与超线性激活函数,探究特征图中的干扰机制。
  • 在150万参数以下模型中,减少干扰使准确率和可扩展性显著提升。
  • 提出NoDepth瓶颈结构,适合资源受限场景下的高效视觉任务。

本文研究当前最先进的低参数深度神经网络在计算机视觉中的表现,聚焦于瓶颈架构及其使用超线性激活函数时的行为。我们关注与表征叠加相关的特征图干扰现象,即神经元同时编码多种特征。研究表明,在极低规模网络(参数少于150万)中限制干扰可提升模型的可扩展性与准确性。通过考察多种瓶颈架构,我们识别出降低干扰的关键设计要素,并据此提出一个基于实验机理洞察的原型架构——NoDepth Bottleneck。该架构在ImageNet数据集上展现出稳健的可扩展性与高准确率。这些发现有助于构建更高效、可扩展的低参数神经网络,并深化了对计算机视觉中瓶颈结构的理解。

原文摘要 · Abstract (English)

The paper investigates the performance of state-of-the-art low-parameter deep neural networks for computer vision, focusing on bottleneck architectures and their behavior using superlinear activation functions. We address interference in feature maps, a phenomenon associated with superposition, where neurons simultaneously encode multiple characteristics. Our research suggests that limiting interference can enhance scaling and accuracy in very low-scaled networks (under 1.5M parameters). We identify key design elements that reduce interference by examining various bottleneck architectures, leading to a more efficient neural network. Consequently, we propose a proof-of-concept architecture named NoDepth Bottleneck built on mechanistic insights from our experiments, demonstrating robust scaling accuracy on the ImageNet dataset. These findings contribute to more efficient and scalable neural networks for the low-parameter range and advance the understanding of bottlenecks in computer vision. https://caiac.pubpub.org/pub/3dh6rsel

低参数模型特征干扰瓶颈结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。