构建首个大规模海浪裂流视频实例分割基准,助力海滩安全监测
RipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and Safety
- 构建184段视频数据集,含21万帧,覆盖全球多国海滩场景
- 引入时序置信度聚合后处理,使关键指标F2提升12.3%
- 专为减少漏检设计,适合海洋安全与计算机视觉交叉研究者
裂流是向海方向流动的强而窄的水流,导致全球大量海滩伤亡事件。由于其形态不规则且缺乏标注数据,准确识别仍具挑战性。为此,我们提出RipVIS,一个专为裂流分割设计的大规模视频实例分割基准。该数据集规模达前人数据集的十倍,包含184段视频(212,328帧),其中150段(163,528帧)含裂流,数据来源涵盖无人机、手机及固定摄像机。数据涵盖波浪破碎、沉积物流动、水色变化等多样视觉场景,覆盖美国、墨西哥、哥斯达黎加、葡萄牙、意大利、希腊、罗马尼亚、斯里兰卡、澳大利亚和新西兰等地。多数视频以5帧/秒标注,确保动态场景精度,另补充34段无裂流视频(48,800帧)。我们在Mask R-CNN、Cascade Mask R-CNN、SparseInst和YOLO11上进行实验,并针对召回率优化采用F2评分。为提升分割性能,提出基于时序置信度聚合(TCA)的新后处理方法。我们提供基准网站(https://ripvis.ai)共享数据、模型与结果,推动社区协作与持续贡献。
原文摘要 · Abstract (English)
Rip currents are strong, localized and narrow currents of water that flow outwards into the sea, causing numerous beach-related injuries and fatalities worldwide. Accurate identification of rip currents remains challenging due to their amorphous nature and the lack of annotated data, which often requires expert knowledge. To address these issues, we present RipVIS, a large-scale video instance segmentation benchmark explicitly designed for rip current segmentation. RipVIS is an order of magnitude larger than previous datasets, featuring $184$ videos ($212,328$ frames), of which $150$ videos ($163,528$ frames) are with rip currents, collected from various sources, including drones, mobile phones, and fixed beach cameras. Our dataset encompasses diverse visual contexts, such as wave-breaking patterns, sediment flows, and water color variations, across multiple global locations, including USA, Mexico, Costa Rica, Portugal, Italy, Greece, Romania, Sri Lanka, Australia and New Zealand. Most videos are annotated at $5$ FPS to ensure accuracy in dynamic scenarios, supplemented by an additional $34$ videos ($48,800$ frames) without rip currents. We conduct comprehensive experiments with Mask R-CNN, Cascade Mask R-CNN, SparseInst and YOLO11, fine-tuning these models for the task of rip current segmentation. Results are reported in terms of multiple metrics, with a particular focus on the $F_2$ score to prioritize recall and reduce false negatives. To enhance segmentation performance, we introduce a novel post-processing step based on Temporal Confidence Aggregation (TCA). RipVIS aims to set a new standard for rip current segmentation, contributing towards safer beach environments. We offer a benchmark website to share data, models, and results with the research community, encouraging ongoing collaboration and future contributions, at https://ripvis.ai.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。