用光流融合与不确定性建模提升烟雾检测可靠性
Reliable Smoke Detection via Optical Flow-Guided Feature Fusion and Transformer-Based Uncertainty Modeling
- 通过光流编码与高斯混合模型融合视觉与运动特征
- 在多个数据集上达到优于现有方法的检测准确率与鲁棒性
- 适合需要高可信度烟雾检测的安防与工业场景
火灾威胁人类生命与基础设施安全,亟需高精度早期预警系统以识别烟雾等燃烧前兆。然而烟雾具有受光照变化、运动动态与环境噪声影响的复杂时空特性,导致传统检测器可靠性下降。为避免多传感器阵列的部署复杂性,本文提出一种基于单目图像的烟雾特征信息融合框架。首先,采用受四色定理启发的双阶段水平集分数阶变分模型进行光流估计,有效保持运动不连续性;随后,将颜色编码的光流图与外观特征通过高斯混合模型融合,生成烟雾区域二值分割掩码。这些融合特征输入新型移位窗口变换器(Shifted-Windows Transformer),其配备多尺度不确定性估计头,并在两阶段学习策略下训练:第一阶段优化检测准确率,第二阶段联合建模偶然不确定性与认知不确定性,学习预测置信度。大量实验表明,该方法在多种评估指标下均显著优于现有先进方法,具备优异泛化能力与鲁棒性,适用于监控、工业安全与自主监测等早期火灾检测场景。
原文摘要 · Abstract (English)
Fire outbreaks pose critical threats to human life and infrastructure, necessitating high-fidelity early-warning systems that detect combustion precursors such as smoke. However, smoke plumes exhibit complex spatiotemporal dynamics influenced by illumination variability, flow kinematics, and environmental noise, undermining the reliability of traditional detectors. To address these challenges without the logistical complexity of multi-sensor arrays, we propose an information-fusion framework by integrating smoke feature representations extracted from monocular imagery. Specifically, a Two-Phase Uncertainty-Aware Shifted Windows Transformer for robust and reliable smoke detection, leveraging a novel smoke segmentation dataset, constructed via optical flow-based motion encoding, is proposed. The optical flow estimation is performed with a four-color-theorem-inspired dual-phase level-set fractional-order variational model, which preserves motion discontinuities. The resulting color-encoded optical flow maps are fused with appearance cues via a Gaussian Mixture Model to generate binary segmentation masks of the smoke regions. These fused representations are fed into the novel Shifted-Windows Transformer, which is augmented with a multi-scale uncertainty estimation head and trained under a two-phase learning regimen. First learning phase optimizes smoke detection accuracy, while during the second phase, the model learns to estimate plausibility confidence in its predictions by jointly modeling aleatoric and epistemic uncertainties. Extensive experiments using multiple evaluation metrics and comparative analysis with state-of-the-art approaches demonstrate superior generalization and robustness, offering a reliable solution for early fire detection in surveillance, industrial safety, and autonomous monitoring applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。