用视觉语言模型比对视频描述,自动找专家操作和决策场景。
Anomalous Frame Detection Using VLM-Based Description Comparison for Extracting Expert-Specific Actions and Contextual Decision-Making Scenes with Intra-Video Self-Similarity
- 用VLM生成帧级描述,比对双视频差异提取专家动作。
- 在27个场景中动作候选提取率达65%,决策场景达61%。
- 适合需要传承专家经验的工业维护领域应用。
铁路、电厂等关键基础设施的维护关乎安全与可靠,但熟练工人短缺促使知识传承成为迫切需求。以往研究多通过对比普通操作与专家操作视频中的可观察动作来提取专家知识,但专家经验不仅体现在动作上,更嵌入任务执行中的情境化决策。本文提出一种基于视觉语言模型(VLM)描述比对的方法,检测两段任务视频间的异常帧,以自动提取包含专家特有动作及情境决策的候选场景。方法通过比较帧级描述的相似性识别专家动作,利用描述的视频内自相似性识别决策场景。在模拟配电盘维护实验中,共涵盖27个任务场景,所提方法对动作候选的提取率达到65%,对决策场景候选的提取率为61%,优于传统方法的59%和33%。结果表明该方法能有效发现蕴含专家经验的候选场景。
原文摘要 · Abstract (English)
Maintenance of critical infrastructures, such as railways and power plants, is essential for ensuring operational safety and reliability. However, the declining number of skilled maintenance workers highlights the need to transfer expert know-how to less experienced workers. Previous studies have attempted to extract candidates of expert knowledge by comparing videos of manual-based work with those of expert workers, mainly focusing on differences in observable actions. However, expert know-how is often embedded not only in actions but also in contextual decision-making during task execution. This paper proposes a method that detects anomalous frames between two task videos to automatically extract candidate scenes containing expert-specific actions and contextual decision-making scenes. The method generates frame-wise visual descriptions using a vision-language model (VLM). Expert-specific actions are extracted based on frame similarities computed from description comparisons between two videos, while contextual decision-making scenes are extracted using segment similarities derived from intra-video self-similarity of the descriptions. In simulated distribution board maintenance experiments involving 27 task scenarios, the proposed method achieved extraction rates of 65% for action candidates and 61% for decision-scene candidates, improving over conventional methods that achieved 59% and 33%, respectively. These results demonstrate the effectiveness of the proposed approach in discovering candidate scenes containing expert know-how.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。