arXiv:2502.12415cs.CV2025-02TPAMI被引 5

首次探索气体目标检测,构建了14万帧视频数据集并提出物理启发的检测框架。

Gaseous Object Detection

  • 基于高斯扩散模型设计可变形状建模的体素位移场(VSF)
  • 在600个视频、14.1万帧上建立基准,实现气体目标检测
  • 为非刚性、无边界的气体检测提供新思路,适合视觉与物理交叉研究者

目标检测是计算机视觉中的基础且具有挑战性的问题,得益于深度学习的高效性得到了快速发展。当前主要针对具有明显视觉特征的刚性固体物体。本文致力于一个极少被探索的任务——气体目标检测(Gaseous Object Detection, GOD),旨在探讨目标检测技术能否从固体扩展到气体。然而,气体具有显著不同的视觉特性:1)显著性不足,2)形状任意且持续变化,3)缺乏明确边界。为推动该挑战性任务的研究,我们构建了GOD-Video数据集,包含600个视频(共141,017帧),涵盖多种气体属性和类型。基于该数据集建立了全面的基准,支持对帧级与视频级检测器的严格评估。受高斯扩散模型启发,设计了物理驱动的体素位移场(VSF),用于建模潜在三维空间中几何不规则性和动态形状变化。将VSF融入Faster RCNN,形成简单而强大的基线模型VSF RCNN。本工作旨在吸引更多研究关注这一有价值但极具挑战的领域。

原文摘要 · Abstract (English)

Object detection, a fundamental and challenging problem in computer vision, has experienced rapid development due to the effectiveness of deep learning. The current objects to be detected are mostly rigid solid substances with apparent and distinct visual characteristics. In this paper, we endeavor on a scarcely explored task named Gaseous Object Detection (GOD), which is undertaken to explore whether the object detection techniques can be extended from solid substances to gaseous substances. Nevertheless, the gas exhibits significantly different visual characteristics: 1) saliency deficiency, 2) arbitrary and ever-changing shapes, 3) lack of distinct boundaries. To facilitate the study on this challenging task, we construct a GOD-Video dataset comprising 600 videos (141,017 frames) that cover various attributes with multiple types of gases. A comprehensive benchmark is established based on this dataset, allowing for a rigorous evaluation of frame-level and video-level detectors. Deduced from the Gaussian dispersion model, the physics-inspired Voxel Shift Field (VSF) is designed to model geometric irregularities and ever-changing shapes in potential 3D space. By integrating VSF into Faster RCNN, the VSF RCNN serves as a simple but strong baseline for gaseous object detection. Our work aims to attract further research into this valuable albeit challenging area.

目标检测气体检测物理建模视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。