用视频和语音联合分析专家操作,自动生成可指导工人的结构化任务知识。
AI-based worker guidance in assembly and disassembly operations using multimodal ego/exo-centric data capture and structured task knowledge
- 融合第一人称与第三人称视频及语音,联合编码时间与多模态信息
- 视频表示比静态图像更准确捕捉流程结构与执行上下文
- 适合维修培训、循环制造等需要知识传承的工业场景
装配与拆解过程依赖难以记录、复用和传递的专家经验。本文提出一种以数据为中心的方法,通过第一人称和第三人称视频记录,从专家示范中提取结构化任务知识。视频与语音中的时序及多模态信息被联合编码,生成可支持流程文档化与情境感知工人引导的结构化任务表征。在真实世界拆解案例研究中评估表明,基于视频的表示能捕捉静态图像方法无法体现的流程结构与执行上下文。结果凸显了第一人称视频理解在维修、培训与循环制造中的潜力。
原文摘要 · Abstract (English)
Assembly and disassembly processes rely on expert knowledge that is difficult to document, reuse, and transfer. This paper presents a data-centric approach for extracting structured task knowledge from expert demonstrations using egocentric and exocentric recordings. Temporal and multimodal information from video and narration is jointly encoded to derive structured task representations that enable procedural documentation and context-aware worker guidance. The approach is evaluated on a real-world disassembly case study, demonstrating that video-based representations capture procedural structure and execution context beyond static image-based methods. The results highlight the potential of egocentric video understanding for repair, training, and circular manufacturing applications. Project website: https://indego-assistant.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。