用视觉语言模型比对视频,自动发现维修专家的隐性操作技巧
Anomalous Frame Detection by Grouping Frame Similarities between Two Videos Computed by Vision-Language Model to Extract Expert Workers' Unique Actions
- 通过对比新手与专家维修视频的帧相似性,定位异常动作
- 在11类手册未记载动作中,知识提取率提升至66.9%,高出传统方法50个百分点
- 适合需要挖掘隐性经验的工业维护场景,尤其适用于技能传承
关键基础设施如铁路和电厂的维护对运行安全至关重要,但熟练维修人员数量减少,使经验传承面临挑战。传统访谈法难以捕捉专家自身也未必意识到的隐性技能。为此,本文提出一种方法:通过比较手动操作视频与专家操作视频的帧间相似性,识别包含隐性知识的异常动作帧。在配电盘模拟维护实验中,针对11类手册未说明的操作,该方法实现66.9%的知识提取率,较传统技术提升50个百分点。结果表明,该方法能有效揭示隐藏的维护知识,助力关键基础设施维护中的技能传递与人才培养。
原文摘要 · Abstract (English)
Maintenance of critical infrastructures, such as railways and power plants, is essential for operational safety and reliability. However, the declining number of skilled maintenance workers poses a serious challenge to sustaining these operations, highlighting the need to effectively transfer expert know-how to less experienced workers. Although traditional interview-based approaches have been used to elicit maintenance skills, they struggle to capture know-how that experts themselves may not consciously recognize. To address this gap, we proposed a method that detects anomalous frames of candidate actions including know-how by comparing a video of manual-based work with that of expert maintenance workers. In a simulated maintenance experiment involving a distribution board, our method targeted 11 types of actions not described in the manual and achieved a 66.9% extraction rate, marking a 50-percentage-point improvement over conventional techniques. These findings underscore the effectiveness of our approach in revealing hidden maintenance knowledge, thereby contributing to enhanced skill transfer and workforce development in critical infrastructure maintenance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。