arXiv:2502.13859cs.CV2025-02被引 3

构建首个多场景大尺度视频伪装目标检测数据集,推动该技术在安防等领域的应用。

MSVCOD:A Large-Scale Multi-Scene Dataset for Video Camouflage Object Detection

  • 设计半自动迭代标注流程,高效生成高质量标注数据。
  • 包含人类、车辆、医疗等新类物体,覆盖多种复杂背景环境。
  • 提出单流模型实现端到端检测,性能超越现有方法。

视频伪装目标检测(VCOD)旨在识别在视频中与背景无缝融合的隐蔽物体。视频的动态特性使运动线索或视角变化成为检测依据。以往的VCOD数据集主要聚焦动物对象,研究范围局限于野生动物场景。然而,VCOD在安全、艺术和医疗等领域具有广泛应用价值。为此,我们构建了首个大规模多领域视频伪装目标检测数据集MSVCOD。为保证标注质量,设计了一套半自动迭代标注流程,在降低人工成本的同时保持高精度。该数据集是目前最大的VCOD数据集,首次引入人体、动物、医疗及车辆等多类目标,并扩展了多种环境背景。这一拓展显著提升了任务的实际应用潜力。同时,我们提出一种单流视频伪装目标检测模型,无需额外运动特征融合模块即可完成特征提取与信息融合。该框架在现有动物类VCOD数据集及新提出的MSVCOD上均达到当前最优性能。数据集与代码将公开发布。

原文摘要 · Abstract (English)

Video Camouflaged Object Detection (VCOD) is a challenging task which aims to identify objects that seamlessly concealed within the background in videos. The dynamic properties of video enable detection of camouflaged objects through motion cues or varied perspectives. Previous VCOD datasets primarily contain animal objects, limiting the scope of research to wildlife scenarios. However, the applications of VCOD extend beyond wildlife and have significant implications in security, art, and medical fields. Addressing this problem, we construct a new large-scale multi-domain VCOD dataset MSVCOD. To achieve high-quality annotations, we design a semi-automatic iterative annotation pipeline that reduces costs while maintaining annotation accuracy. Our MSVCOD is the largest VCOD dataset to date, introducing multiple object categories including human, animal, medical, and vehicle objects for the first time, while also expanding background diversity across various environments. This expanded scope increases the practical applicability of the VCOD task in camouflaged object detection. Alongside this dataset, we introduce a one-steam video camouflage object detection model that performs both feature extraction and information fusion without additional motion feature fusion modules. Our framework achieves state-of-the-art results on the existing VCOD animal dataset and the proposed MSVCOD. The dataset and code will be made publicly available.

视频检测伪装目标数据集多场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。