用视觉语言模型检测实验流程中的异常,提升科研机器人安全性和稳定性。
A VLM-based Method for Visual Anomaly Detection in Robotic Scientific Laboratories
- 基于视觉语言模型,通过四种提示配置实现多层级监督。
- 上下文信息越丰富,检测准确率越高,验证了方法有效性。
- 适用于科研实验室自动化流程的异常监测,适合实验安全研究者。
在机器人科学实验室中,视觉异常检测对于及时发现和解决潜在故障或偏差至关重要,已成为保障实验过程稳定与安全的关键因素。为应对这一挑战,本文提出一种基于视觉语言模型(VLM)的视觉推理方法,通过四种逐步提供更多信息的提示配置,支持不同层级的监督。为系统评估其有效性,我们构建了一个针对科学工作流过程异常检测的视觉基准数据集。在两种代表性视觉语言模型上的实验表明,随着上下文信息的增加,检测准确率提升,证实了所提推理方法在科学工作流异常检测中的有效性和适应性。此外,在选定实验步骤的真实世界验证中,第一人称视觉观测可有效识别流程级异常。该工作为科学实验工作流中的视觉异常检测提供了数据驱动基础与评估框架。
原文摘要 · Abstract (English)
In robot scientific laboratories, visual anomaly detection is important for the timely identification and resolution of potential faults or deviations. It has become a key factor in ensuring the stability and safety of experimental processes. To address this challenge, this paper proposes a VLM-based visual reasoning approach that supports different levels of supervision through four progressively informative prompt configurations. To systematically evaluate its effectiveness, we construct a visual benchmark tailored for process anomaly detection in scientific workflows. Experiments on two representative vision-language models show that detection accuracy improves as more contextual information is provided, confirming the effectiveness and adaptability of the proposed reasoning approach for process anomaly detection in scientific workflows. Furthermore, real-world validations at selected experimental steps confirm that first-person visual observation can effectively identify process-level anomalies. This work provides both a data-driven foundation and an evaluation framework for vision anomaly detection in scientific experiment workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。