arXiv:2505.04397cs.CVcs.AI2025-05中稿 · the GCPR 2026被引 2

用乘法机制增强视觉网络,提升性能且不增加参数量

PURe: A Plug-and-Play Product-Unit Residual Module for Vision Networks

  • 引入二维乘积单元,通过对数域设计实现深层网络中的乘法聚合
  • 在多个数据集上以更少参数达到或超过更深残差网络的精度
  • 可直接替换原有残差模块,适合追求高效模型的研究者

现代视觉网络主要依赖加性局部变换,而显式的乘性局部交互仍研究不足。乘积单元可直接建模此类交互,但其在深度架构中的应用受限于优化不稳定性。本文提出PURe,一种用于深度视觉网络的乘积单元残差模块。PURe基于二维乘积单元,采用实值对数域公式,使乘法局部聚合在深层残差层级中变得可行。该模块可作为原生残差单元的即插即用替代品。我们在残差CNN上应用于图像分类,在2D残差编码器-解码器网络上应用于体积分层CT数据的切片级分割。在Galaxy10 DECaLS、ImageNet和CIFAR-10上,PURe持续提升残差CNN性能,并实现更优的准确率-参数权衡,使中等深度模型以更小参数量达到或超越显著更深的ResNet基线。在AMOS基准上,PURe也在3D病例级评估下提升了切片级CT分割表现。结果表明,显式乘法局部交互是深度残差视觉网络中实用且有效的设计基础。

原文摘要 · Abstract (English)

Modern vision networks are dominated by additive local transformations, whereas explicit multiplicative local interactions remain underexplored. Product units offer a direct approach to modeling such interactions, but their use in deep architectures has been limited by optimization instability. In this work, we propose PURe, a product-unit Residual Module for deep vision networks. PURe is built around a 2D product unit with a real-valued log-domain formulation that makes multiplicative local aggregation practical within deep residual hierarchies. The resulting module serves as a drop-in replacement for native residual units. We instantiate PURe in residual CNNs for image classification and in 2D residual encoder--decoder networks for slice-based segmentation on volumetric CT data. Across Galaxy10 DECaLS, ImageNet, and CIFAR-10, PURe consistently improves residual CNNs and yields a more favorable accuracy--parameter trade-off, allowing moderately deep models to match or surpass substantially deeper ResNet baselines with much smaller parameter budgets. On the AMOS benchmark, PURe also improves slice-based CT segmentation under 3D case-level evaluation. These results show that explicit multiplicative local interaction is a practical and effective design primitive for deep residual vision networks.

视觉网络乘积单元残差模块模型效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。