arXiv:2409.07813cs.CV2024-09被引 122

YOLOv9通过新架构与训练方法,实现更准更快的实时目标检测。

What is YOLOv9: An In-Depth Exploration of the Internal Features of the Next-Generation Object Detector

  • 采用GELAN和PGI提升特征提取与梯度流动效率
  • COCO数据集上mAP更高、推理速度更快,优于YOLOv8
  • 适配边缘设备到高性能GPU,部署灵活

本研究全面分析了YOLOv9目标检测模型,聚焦其架构创新、训练方法及相比前代的性能提升。关键改进包括通用高效层聚合网络(GELAN)与可编程梯度信息(PGI),显著增强特征提取与梯度流,提升准确率与效率。通过引入深度可分离卷积与轻量级C3Ghost结构,降低计算复杂度同时保持高精度。在Microsoft COCO基准测试中,YOLOv9在均值平均精度(mAP)和推理速度上均超越YOLOv8。模型具备跨平台部署能力,支持从边缘设备到高性能GPU的无缝运行,并原生集成PyTorch与TensorRT。本文首次深入解析YOLOv9内部特性及其实际应用潜力,确立其作为工业级实时目标检测前沿解决方案的地位。

原文摘要 · Abstract (English)

This study provides a comprehensive analysis of the YOLOv9 object detection model, focusing on its architectural innovations, training methodologies, and performance improvements over its predecessors. Key advancements, such as the Generalized Efficient Layer Aggregation Network GELAN and Programmable Gradient Information PGI, significantly enhance feature extraction and gradient flow, leading to improved accuracy and efficiency. By incorporating Depthwise Convolutions and the lightweight C3Ghost architecture, YOLOv9 reduces computational complexity while maintaining high precision. Benchmark tests on Microsoft COCO demonstrate its superior mean Average Precision mAP and faster inference times, outperforming YOLOv8 across multiple metrics. The model versatility is highlighted by its seamless deployment across various hardware platforms, from edge devices to high performance GPUs, with built in support for PyTorch and TensorRT integration. This paper provides the first in depth exploration of YOLOv9s internal features and their real world applicability, establishing it as a state of the art solution for real time object detection across industries, from IoT devices to large scale industrial applications.

目标检测YOLOv9实时系统轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。