arXiv:2601.04727cs.CVcs.NE2026-01

定制轻量CNN在5个异构图像数据集上实现高效跨域分类

Training a Custom CNN on Five Heterogeneous Image Datasets

  • 设计轻量级定制CNN,适配多领域图像特征
  • 在数据有限场景下,迁移学习显著提升模型性能
  • 为资源受限的现实视觉任务提供实用部署方案

深度学习已彻底改变视觉数据分析,卷积神经网络(CNN)能够直接从图像中自动学习有意义的特征表示,相比传统手工特征工程方法更具优势。本研究考察了基于CNN的架构在五个异构数据集上的有效性,涵盖农业与城市领域:芒果品种分类、水稻品种识别、道路表面状况评估、电动三轮车检测和人行道侵占监测。这些数据集存在光照差异、分辨率不一、环境复杂度高及类别不平衡等挑战,要求模型具备适应性和鲁棒性。我们评估了一个轻量级、任务定制的自定义CNN,并对比了经典深度架构如ResNet-18和VGG-16,采用从零训练和迁移学习两种方式。通过系统化的预处理、数据增强与受控实验,分析了模型复杂度、深度及预训练对收敛速度、泛化能力与跨数据集性能的影响。主要贡献包括:(1) 构建了一个在多个应用领域表现优异的高效自定义CNN;(2) 提供全面比较分析,揭示迁移学习与深层架构在数据稀缺环境下带来的显著优势。研究结果为资源受限但高影响力的现实视觉分类任务提供了切实可行的深度学习部署指导。

原文摘要 · Abstract (English)

Deep learning has transformed visual data analysis, with Convolutional Neural Networks (CNNs) becoming highly effective in learning meaningful feature representations directly from images. Unlike traditional manual feature engineering methods, CNNs automatically extract hierarchical visual patterns, enabling strong performance across diverse real-world contexts. This study investigates the effectiveness of CNN-based architectures across five heterogeneous datasets spanning agricultural and urban domains: mango variety classification, paddy variety identification, road surface condition assessment, auto-rickshaw detection, and footpath encroachment monitoring. These datasets introduce varying challenges, including differences in illumination, resolution, environmental complexity, and class imbalance, necessitating adaptable and robust learning models. We evaluate a lightweight, task-specific custom CNN alongside established deep architectures, including ResNet-18 and VGG-16, trained both from scratch and using transfer learning. Through systematic preprocessing, augmentation, and controlled experimentation, we analyze how architectural complexity, model depth, and pre-training influence convergence, generalization, and performance across datasets of differing scale and difficulty. The key contributions of this work are: (1) the development of an efficient custom CNN that achieves competitive performance across multiple application domains, and (2) a comprehensive comparative analysis highlighting when transfer learning and deep architectures provide substantial advantages, particularly in data-constrained environments. These findings offer practical insights for deploying deep learning models in resource-limited yet high-impact real-world visual classification tasks.

CNN跨域分类迁移学习轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。