arXiv:2502.05500cs.ROcs.AI2025-02被引 3

用视觉+超声融合检测工业泄漏与电弧,准确率达99%。

Vision-Ultrasound Robotic System based on Deep Learning for Gas and Arc Hazard Detection in Manufacturing

  • 融合视觉与超声波的深度学习检测框架,实现自主识别。
  • 在噪声环境中检测准确率比传统方法高44个百分点。
  • 适用于工厂巡检机器人,推理仅需2.1秒,实时性好。

气体泄漏和电弧放电在工业环境中构成重大风险,亟需可靠的检测系统以保障安全与运行效率。受人类通过视觉识别结合听觉验证的启发,本研究提出一种基于深度学习的机器人系统,可自主检测并分类制造环境中的气体泄漏与电弧放电。系统所有实验任务均在机器人本地完成。采用112通道超声麦克风阵列,采样率为96 kHz,捕捉超声频段信号,处理在多种工业场景下采集的真实数据集,涵盖不同泄漏类型(如针孔、开放端)及部分放电类型(电晕、表面、悬浮),且在不同噪声条件下测试。系统整合视觉检测与束形成增强的声学分析流程,信号经STFT变换后通过伽马校正优化,实现鲁棒特征提取。采用类Inception结构的CNN进行分类,气体泄漏检测准确率达99%。系统不仅能定位单一危险源,还能通过融合视觉与声学多模态数据提升分类可靠性。在混响与噪声增强环境下,性能较传统模型最高提升44个百分点。实验设计确保公平性与可复现性。系统优化为实时部署,在移动机器人平台上的推理时间仅为2.1秒。通过模仿人类巡检流程,融合多模态感知,该研究为工业自动化提供高效安全解决方案。

原文摘要 · Abstract (English)

Gas leaks and arc discharges present significant risks in industrial environments, requiring robust detection systems to ensure safety and operational efficiency. Inspired by human protocols that combine visual identification with acoustic verification, this study proposes a deep learning-based robotic system for autonomously detecting and classifying gas leaks and arc discharges in manufacturing settings. The system is designed to execute all experimental tasks entirely onboard the robot. Utilizing a 112-channel acoustic camera operating at a 96 kHz sampling rate to capture ultrasonic frequencies, the system processes real-world datasets recorded in diverse industrial scenarios. These datasets include multiple gas leak configurations (e.g., pinhole, open end) and partial discharge types (Corona, Surface, Floating) under varying environmental noise conditions. Proposed system integrates visual detection and a beamforming-enhanced acoustic analysis pipeline. Signals are transformed using STFT and refined through Gamma Correction, enabling robust feature extraction. An Inception-inspired CNN further classifies hazards, achieving 99% gas leak detection accuracy. The system not only detects individual hazard sources but also enhances classification reliability by fusing multi-modal data from both vision and acoustic sensors. When tested in reverberation and noise-augmented environments, the system outperformed conventional models by up to 44%p, with experimental tasks meticulously designed to ensure fairness and reproducibility. Additionally, the system is optimized for real-time deployment, maintaining an inference time of 2.1 seconds on a mobile robotic platform. By emulating human-like inspection protocols and integrating vision with acoustic modalities, this study presents an effective solution for industrial automation, significantly improving safety and operational reliability.

工业安全多模态检测机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。