用视觉+状态机实现热锻工件近实时定位,精度达317.8毫米
Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines
- 以设备为中心,通过多相机视觉识别抓放动作
- 事件驱动状态机实现100%事件检测,平均定位误差317.8毫米
- 适合工业现场追踪与设备效率分析,可解释性强
热锻过程中连续工件定位对可追溯性和流程协调至关重要,但因高温、表面劣化和路径不规则,直接跟踪不可靠。本文提出一种设备中心框架,通过多个固定2D摄像头观测处理设备来推断工件位置。该框架估计地板平面空间中的3D设备坐标,并识别抓取与释放行为。事件驱动的有限状态机将这些行为验证为离散操作事件,并持续更新工件状态与位置。在3D卷积神经网络中引入关键点引导注意力机制,聚焦功能相关设备区域,提升行为识别性能。在实际热锻工厂评估中,系统在33秒容差窗口内实现100%事件检测准确率,平均定位误差为317.8毫米,平均系统延迟21秒。该框架实现了视觉感知与可解释事件驱动推理的结合,支持工件转移可视化及设备操作定量分析。
原文摘要 · Abstract (English)
Continuous workpiece localization is essential for traceability and process coordination in hot forging, but direct tracking is unreliable because of extreme temperatures, surface degradation, and irregular routing. This study presents an equipment-centric framework that infers workpiece locations from handling equipment observed by multiple static 2D cameras. The framework estimates floorplan-space 3D equipment coordinates and recognizes grasp and release activities. Event-driven finite state machines validate these activities as discrete handling events and continuously update workpiece states and locations. A keypoint-guided attention mechanism integrated into a 3D convolutional neural network improves activity recognition by focusing on functionally relevant equipment regions. Evaluation in an operational hot forging factory achieved 100\% event detection accuracy within a 33-second tolerance window, a mean localization error of 317.8 mm, and a mean system latency of 21 seconds. The framework connects vision-based perception with interpretable event-driven reasoning and supports visualization of workpiece transfers and quantitative analysis of equipment operations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。