arXiv:2511.06295cs.CV2025-11被引 1

用单摄像头实现叉车自动识别货盘与卡槽,降低成本提升效率。

Learning-Based Vision Systems for Semi-Autonomous Forklift Operation in Industrial Warehouse Environments

  • 基于YOLOv8/v11和超参优化,单摄像头检测货盘与卡槽。
  • 优化后YOLOv11精度更高,收敛稳定,检测准确率优异。
  • 适合工业场景改造,可低成本部署于现有叉车系统。

仓库物料搬运自动化依赖于鲁棒且低成本的感知系统。本文提出一种基于单个标准摄像头的视觉框架,用于货盘与货盘卡槽的检测与建图。采用YOLOv8与YOLOv11架构,通过Optuna驱动的超参数优化及空间后处理增强性能。创新的货盘卡槽建模模块将检测结果转化为可操作的空间表示,实现货盘与卡槽的精准关联,支持叉车作业。在包含真实仓库图像的自建数据集上测试显示,YOLOv8具备高检测准确率;而经过优化的YOLOv11在精度和收敛稳定性上表现更优。结果证明该方案可行,可作为低成本、可改造的视觉感知模块应用于叉车。本研究为推进仓储自动化提供了可扩展路径,促进更安全、经济、智能的物流运作。

原文摘要 · Abstract (English)

The automation of material handling in warehouses increasingly relies on robust, low cost perception systems for forklifts and Automated Guided Vehicles (AGVs). This work presents a vision based framework for pallet and pallet hole detection and mapping using a single standard camera. We utilized YOLOv8 and YOLOv11 architectures, enhanced through Optuna driven hyperparameter optimization and spatial post processing. An innovative pallet hole mapping module converts the detections into actionable spatial representations, enabling accurate pallet and pallet hole association for forklift operation. Experiments on a custom dataset augmented with real warehouse imagery show that YOLOv8 achieves high pallet and pallet hole detection accuracy, while YOLOv11, particularly under optimized configurations, offers superior precision and stable convergence. The results demonstrate the feasibility of a cost effective, retrofittable visual perception module for forklifts. This study proposes a scalable approach to advancing warehouse automation, promoting safer, economical, and intelligent logistics operations.

视觉感知叉车自动化目标检测工业机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。