融合CNN与视觉大模型,实现高速视频相变分割的高精度与可靠性
MSEG-VCUQ: Multimodal SEGmentation with Enhanced Vision Foundation Models, Convolutional Neural Networks, and Uncertainty Quantification for High-Speed Video Phase Detection Data
- 用U-Net与SAM融合提升多模态视频分割精度
- 首次引入像素级不确定性量化,关键指标误差更可控
- 开源首个高速相变视频数据集,助力工业沸腾研究
高速视频(HSV)相变检测(PD)分割对监测工业过程中气相、液相及微层相至关重要。尽管基于卷积神经网络(CNN)的U-Net在简化阴影成像两相流分析中表现良好,但其在复杂HSV PD任务中的应用仍待探索,而视觉基础模型(VFMs)也尚未解决阴影成像或PD两相流视频分割的复杂性问题。现有不确定性量化(UQ)方法缺乏像素级可靠性,无法准确评估接触线密度和干区占比等关键指标,且缺乏针对PD分割的大规模多模态实验数据集,严重制约了该领域发展。为此,本文提出MSEG-VCUQ:一种融合U-Net CNN与基于Transformer的Segment Anything Model(SAM)的混合框架,显著提升分割精度与跨模态泛化能力;系统引入不确定性量化以实现稳健误差评估,并构建首个开源的多模态高速视频相变检测数据集。实验表明,MSEG-VCUQ优于基线CNN与视觉基础模型,可实现可扩展、可靠的实时沸腾动力学分割。
原文摘要 · Abstract (English)
High-speed video (HSV) phase detection (PD) segmentation is crucial for monitoring vapor, liquid, and microlayer phases in industrial processes. While CNN-based models like U-Net have shown success in simplified shadowgraphy-based two-phase flow (TPF) analysis, their application to complex HSV PD tasks remains unexplored, and vision foundation models (VFMs) have yet to address the complexities of either shadowgraphy-based or PD TPF video segmentation. Existing uncertainty quantification (UQ) methods lack pixel-level reliability for critical metrics like contact line density and dry area fraction, and the absence of large-scale, multimodal experimental datasets tailored to PD segmentation further impedes progress. To address these gaps, we propose MSEG-VCUQ. This hybrid framework integrates U-Net CNNs with the transformer-based Segment Anything Model (SAM) to achieve enhanced segmentation accuracy and cross-modality generalization. Our approach incorporates systematic UQ for robust error assessment and introduces the first open-source multimodal HSV PD datasets. Empirical results demonstrate that MSEG-VCUQ outperforms baseline CNNs and VFMs, enabling scalable and reliable PD segmentation for real-world boiling dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。