构建带像素级标注的水下视频数据集,解决动态环境下的视觉分析难题。
AUTV: Creating Underwater Video Datasets with Pixel-wise Annotations

- 提出AUTV框架,合成带像素标注的水下视频数据。
- 建成两个数据集:UTV含2000对真实视频与文本,SUTV含1万条带分割掩码的合成视频。
- 适用于水下视频修复与目标分割等下游任务,提升模型性能。
水下视频分析因动态海洋环境和相机运动而面临挑战。现有无训练视频生成方法基于逐帧学习运动规律,常导致明显运动中断与错位。为此,我们提出AUTV框架,用于合成带像素级标注的海洋视频数据。通过该框架构建了两个视频数据集:UTV为真实世界数据集,包含2000个视频-文本对;SUTV为合成视频数据集,含10000条带有海洋物体分割掩码的视频。UTV提供多样化的水下视频,涵盖外观、纹理、相机内参、光照及动物行为等丰富标注。SUTV可用于提升水下下游任务性能,已在视频修复和视频目标分割中得到验证。
原文摘要 · Abstract (English)
Underwater video analysis, hampered by the dynamic marine environment and camera motion, remains a challenging task in computer vision. Existing training-free video generation techniques, learning motion dynamics on the frame-by-frame basis, often produce poor results with noticeable motion interruptions and misaligments. To address these issues, we propose AUTV, a framework for synthesizing marine video data with pixel-wise annotations. We demonstrate the effectiveness of this framework by constructing two video datasets, namely UTV, a real-world dataset comprising 2,000 video-text pairs, and SUTV, a synthetic video dataset including 10,000 videos with segmentation masks for marine objects. UTV provides diverse underwater videos with comprehensive annotations including appearance, texture, camera intrinsics, lighting, and animal behavior. SUTV can be used to improve underwater downstream tasks, which are demonstrated in video inpainting and video object segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。