用视觉与语言联合模型提升工业气体泄漏的精准分割
JVLGS: Joint Vision-Language Gas Leak Segmentation
- 融合图像与文本信息,增强泄漏区域识别能力
- 在多种场景下显著优于现有方法,少样本学习也表现稳定
- 适配真实工业环境,对误报有强抑制能力
气体泄漏对人类健康和工业安全构成严重威胁。然而,实现准确及时的泄漏监测仍是重大挑战。现有基于红外图像的视觉方法受限于泄漏羽流固有的模糊性和非刚性特征,导致检测可靠性与精度下降。为此,本文提出一种联合视觉-语言气体泄漏分割框架(JVLGS),通过整合视觉与文本模态的互补优势,提升泄漏分割性能。考虑到泄漏事件偶发,视频中大量帧无泄漏,JVLGS引入自适应后处理模块,有效抑制由噪声和非目标物体引发的误报——这是现有方法的常见短板。在多样工业场景下的大量实验表明,JVLGS显著优于当前最先进方法。此外,其在监督学习与小样本学习设置下均表现稳健,而对比方法通常仅在单一设置下表现良好或在两者中均表现不佳。
原文摘要 · Abstract (English)
Gas leaks pose severe risks to human health and industrial safety. However, accurate and timely monitoring of gas leaks remains a major challenge. Existing vision-based methods using infrared (IR) imagery are limited by the inherently blurry and non-rigid nature of leak plumes, which reduces detection reliability and precision. To overcome these limitations, this paper proposes a Joint Vision-Language Gas leak Segmentation (JVLGS) framework that integrates the complementary strengths of visual and textual modalities to enhance gas leak segmentation. Recognizing that gas leaks are sporadic and many video frames contain no leakage, JVLGS incorporates an adaptive postprocessing module to effectively suppress false positives caused by noise and non-target objects-a common limitation of existing approaches. Extensive experiments across diverse industrial scenarios demonstrate that JVLGS significantly outperforms state-of-the-art gas leak segmentation methods. Furthermore, it achieves consistently strong performance under both supervised and few-shot learning settings, whereas competing methods typically perform well in only one setting or underperform in both.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。