用AI检测海上船只,ViT模型表现最佳
AI for Maritime Security: Comparative Evaluation of CNN and Vision Transformer Architectures for Maritime Object Detection
- 对比六种深度模型,用6468张海面图像测试
- ViT达100%准确率,错误率最低且处理最快
- 适合海上监控、边境防卫等实时场景
本研究旨在通过先进的人工智能与计算机视觉技术提升海上安全。为此,设计并评估了可在不同实时环境下检测海面船只的智能目标检测系统。基于包含6,468张图像的海事图像数据集,覆盖阴天、雾天、雨天和晴天等多种天气条件,评估了六种深度学习架构:一个基础卷积神经网络(CNN)模型、四种迁移学习模型(Xception、VGG16、MobileNetV2、EfficientNetV2L)以及一个视觉变换器(ViT)模型。通过准确率、Ⅰ型与Ⅱ型错误率、模型大小和视频处理时间等多项指标进行比较。结果表明,模型性能受计算资源和部署条件影响显著;轻量级架构适用于资源受限设备,而ViT在整体表现上最优,达到100%准确率,错误率最低,视频处理速度最快。研究凸显了基于AI的计算机视觉系统在海上监视、边界保护和自主导航中的潜力。
原文摘要 · Abstract (English)
This study aims to enhance maritime security by using advanced Artificial Intelligence (AI) and Computer Vision (CV) techniques. For this purpose, it was designed and assessed intelligent object detection systems that can detect the presence of ships on the sea surface under different real-time environments. To achieve this goal, a maritime image dataset with 6,468 images was used, covering different weather conditions like cloudy, foggy, rainy, and sunny environments. Six deep learning architectures were evaluated, including a base Convolutional Neural Network (CNN) model, four transfer learning models (Xception, VGG16, MobileNetV2, and EfficientNetV2L), and a Vision Transformer (ViT) model. The models were compared using multiple performance indicators, including accuracy, Type I and Type II errors, model size, and video processing time. The results show that model performance varies depending on computational constraints and deployment conditions. While lightweight architectures are suitable for resource-limited devices, the ViT achieved the best overall performance, reaching 100% accuracy with the lowest error rates and the fastest video processing time. The findings highlight the potential of AI-driven computer vision systems for maritime surveillance, border protection, and autonomous navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。