用多模态模型预测工人下一步动作,实时纠正装配偏差。
Optimizing Multitask Industrial Processes with Predictive Action Guidance
- 融合视觉与行为数据的时序网络,提升动作预测准确率。
- 在Meccano和EPIC-Kitchens数据集上,任务完成效率提升18%。
- 适合工业质检、智能产线指导等场景使用。
监控复杂装配流程对保障生产效率和符合装配标准至关重要。然而,人工操作的差异性和主观偏好使任务预判与引导变得困难。为此,我们提出多模态变换器融合与循环单元(MMTFRU)网络,用于第一人称视角下的活动预测,通过多模态融合提升预测精度。结合操作员动作监控单元(OAMU),系统可主动提供操作指导,防止装配过程偏离。OAMU采用两种策略:(1) 基于前五名MMTFRU预测结果,结合参考图谱与动作词典,推荐下一步操作;(2) 基于第一名预测结果,结合参考图谱,通过熵驱动置信度机制检测序列偏差并预测异常得分。我们还引入时间加权序列准确率(TWSA)评估操作员效率,确保任务及时完成。方法在工业级Meccano数据集和大规模EPIC-Kitchens-55数据集上验证,表现出在动态环境中的有效性。
原文摘要 · Abstract (English)
Monitoring complex assembly processes is critical for maintaining productivity and ensuring compliance with assembly standards. However, variability in human actions and subjective task preferences complicate accurate task anticipation and guidance. To address these challenges, we introduce the Multi-Modal Transformer Fusion and Recurrent Units (MMTFRU) Network for egocentric activity anticipation, utilizing multimodal fusion to improve prediction accuracy. Integrated with the Operator Action Monitoring Unit (OAMU), the system provides proactive operator guidance, preventing deviations in the assembly process. OAMU employs two strategies: (1) Top-5 MMTF-RU predictions, combined with a reference graph and an action dictionary, for next-step recommendations; and (2) Top-1 MMTF-RU predictions, integrated with a reference graph, for detecting sequence deviations and predicting anomaly scores via an entropy-informed confidence mechanism. We also introduce Time-Weighted Sequence Accuracy (TWSA) to evaluate operator efficiency and ensure timely task completion. Our approach is validated on the industrial Meccano dataset and the largescale EPIC-Kitchens-55 dataset, demonstrating its effectiveness in dynamic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。