arXiv:2410.21304cs.CVcs.LG2024-10中稿 · IEEE SSD 2025被引 1

VideoSAM通过微调SAM模型,实现高速视频中复杂气泡的精准分割。

VideoSAM: A Large Vision Foundation Model for High-Speed Video Segmentation

  • 基于SAM架构,针对高速视频数据微调,提升相态检测能力。
  • 在水、氟化液、氮气、氩气四种流体中均显著优于U-Net。
  • 开源了首个面向相态检测的高速视频分割数据集,适合工业与科研应用。

高速视频(HSV)分割对科学与工业中动态物理过程分析至关重要,如沸腾传热。现有模型如U-Net在泛化性和复杂气泡形态分割上表现不足。本文提出VideoSAM,一种针对高速视频相态检测优化的Segment Anything Model(SAM)适配模型。通过多组实验,VideoSAM在水、FC-72、氮气和氩气四种流体环境中均表现出色,显著优于U-Net。此外,我们开源了一个用于相态检测的高速视频分割数据集,推动该领域研究发展。结果表明,VideoSAM有望成为高效、高精度高速视频分割的新标准。代码与数据集已公开于 https://github.com/chikap421/videosam。

原文摘要 · Abstract (English)

High-speed video (HSV) segmentation is essential for analyzing dynamic physical processes in scientific and industrial applications, such as boiling heat transfer. Existing models like U-Net struggle with generalization and accurately segmenting complex bubble formations. We present VideoSAM, a specialized adaptation of the Segment Anything Model (SAM), fine-tuned on a diverse HSV dataset for phase detection. Through diverse experiments, VideoSAM demonstrates superior performance across four fluid environments -- Water, FC-72, Nitrogen, and Argon -- significantly outperforming U-Net in complex segmentation tasks. In addition to introducing VideoSAM, we contribute an open-source HSV segmentation dataset designed for phase detection, enabling future research in this domain. Our findings underscore VideoSAM's potential to set new standards in robust and accurate HSV segmentation. The code and dataset used in this study are available online at https://github.com/chikap421/videosam.

视频分割SAM相态检测高速视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。