arXiv:2503.16546cs.CVcs.AI2025-03综述被引 22

系统梳理深度CNN架构十年演进,覆盖模型创新与应用前沿

A Comprehensive Survey on Architectural Advances in Deep CNNs: Challenges, Applications, and Emerging Research Directions

  • 按空间、路径、通道等维度构建统一分类框架
  • 涵盖2015-2025年关键结构改进与多领域应用成果
  • 适合关注CV与AI架构演进的研究者与工程师

深度卷积神经网络(CNN)在计算机视觉、自然语言处理、医疗诊断、目标检测和语音识别等领域推动了深度学习的突破。架构创新包括一维、二维、三维卷积,空洞卷积、分组卷积、深度可分离卷积及注意力机制,有效应对特定任务挑战,提升特征表示能力与计算效率。结构优化如空间-通道联合建模、多路径设计、特征图增强,强化了层次化特征提取与泛化能力,尤其在迁移学习中表现优异。高效预处理策略如傅里叶变换、结构化变换、低精度计算与权重压缩,显著提升推理速度,支持资源受限环境部署。本文提出统一分类体系,涵盖空间利用、多路径结构、深度、宽度、维度扩展、通道增强与注意力机制。系统回顾了人脸识别、姿态估计、动作识别、文本分类、统计语言建模、疾病诊断、放射学分析、加密货币情绪预测、一维数据处理、视频分析与语音识别等应用。除整合架构进展外,还强调少样本、零样本、弱监督、联邦学习等新兴学习范式。未来研究方向包括混合CNN-Transformer模型、视觉-语言融合、生成学习等。本综述全面呈现2015至2025年CNN的发展脉络,总结关键创新、挑战与机遇。

原文摘要 · Abstract (English)

Deep Convolutional Neural Networks (CNNs) have significantly advanced deep learning, driving breakthroughs in computer vision, natural language processing, medical diagnosis, object detection, and speech recognition. Architectural innovations including 1D, 2D, and 3D convolutional models, dilated and grouped convolutions, depthwise separable convolutions, and attention mechanisms address domain-specific challenges and enhance feature representation and computational efficiency. Structural refinements such as spatial-channel exploitation, multi-path design, and feature-map enhancement contribute to robust hierarchical feature extraction and improved generalization, particularly through transfer learning. Efficient preprocessing strategies, including Fourier transforms, structured transforms, low-precision computation, and weight compression, optimize inference speed and facilitate deployment in resource-constrained environments. This survey presents a unified taxonomy that classifies CNN architectures based on spatial exploitation, multi-path structures, depth, width, dimensionality expansion, channel boosting, and attention mechanisms. It systematically reviews CNN applications in face recognition, pose estimation, action recognition, text classification, statistical language modeling, disease diagnosis, radiological analysis, cryptocurrency sentiment prediction, 1D data processing, video analysis, and speech recognition. In addition to consolidating architectural advancements, the review highlights emerging learning paradigms such as few-shot, zero-shot, weakly supervised, federated learning frameworks and future research directions include hybrid CNN-transformer models, vision-language integration, generative learning, etc. This review provides a comprehensive perspective on CNN's evolution from 2015 to 2025, outlining key innovations, challenges, and opportunities.

深度学习卷积神经网络架构演进综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。