无需标记或固定布局,用AI实时监控工业装配流程
AI-driven visual monitoring of industrial assembly tasks
- 多视角视频+先验知识推理,实现无约束环境下的动作识别
- 在乐高组件更换和液压模具重组任务中准确率超90%
- 适合智能制造、工厂安全监控场景的工程师与研究者
工业装配过程的视觉监控对防止因操作错误导致设备损坏和保障工人安全至关重要。尽管已有商用方案,但通常需要固定的作业空间或贴附视觉标记来简化问题。我们提出ViMAT,一种无需此类限制的AI驱动实时视觉监控系统。该系统由感知模块(从多视角视频流提取视觉信息)与推理模块(基于当前装配状态和先验任务知识推断最可能执行的动作)构成。我们在两个装配任务上验证了ViMAT的有效性:乐高组件更换与液压压机模具重配置。在具有部分遮挡和不确定视觉观测的复杂真实场景中,通过定量与定性分析均展现出优异性能。
原文摘要 · Abstract (English)
Visual monitoring of industrial assembly tasks is critical for preventing equipment damage due to procedural errors and ensuring worker safety. Although commercial solutions exist, they typically require rigid workspace setups or the application of visual markers to simplify the problem. We introduce ViMAT, a novel AI-driven system for real-time visual monitoring of assembly tasks that operates without these constraints. ViMAT combines a perception module that extracts visual observations from multi-view video streams with a reasoning module that infers the most likely action being performed based on the observed assembly state and prior task knowledge. We validate ViMAT on two assembly tasks, involving the replacement of LEGO components and the reconfiguration of hydraulic press molds, demonstrating its effectiveness through quantitative and qualitative analysis in challenging real-world scenarios characterized by partial and uncertain visual observations. Project page: https://tev-fbk.github.io/ViMAT
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。