提出隐蔽密集预测任务,聚焦目标隐藏场景的精准感知。
Deep Learning in Concealed Dense Prediction
- 构建隐蔽对抗分类体系,系统梳理25种SOTA方法
- 在12个隐蔽数据集上对比实验,揭示细粒度表征关键性
- 适合关注农业/工业视觉感知的开发者与研究者
深度学习快速发展,在常规计算机视觉任务中表现优异。随着模型规模、知识容量和推理能力持续提升,是时候关注更复杂的视觉任务了。本文引入并综述一类复杂任务——隐蔽密集预测(Concealed Dense Prediction, CDP),其在农业、工业等领域具有重要价值。CDP的核心特征是目标被环境完全遮蔽,因此需要细粒度表征、先验知识与辅助推理才能充分感知。本文贡献有三方面:(i) 阐明CDP的任务范畴、特性与挑战,强调其与通用视觉任务的本质差异;(ii) 基于隐蔽对抗机制构建分类体系,通过在三个任务上的实验,对比分析25种先进方法在12个常用隐蔽数据集上的表现;(iii) 探讨大模型时代CDP的潜在应用,并总结6个未来研究方向。我们还构建了大规模多模态指令微调数据集CvpINST和隐蔽视觉感知代理CvpAgent,为后续发展提供新视角。
原文摘要 · Abstract (English)
Deep learning is developing rapidly and handling common computer vision tasks well. It is time to pay attention to more complex vision tasks, as model size, knowledge, and reasoning capabilities continue to improve. In this paper, we introduce and review a family of complex tasks, termed Concealed Dense Prediction (CDP), which has great value in agriculture, industry, etc. CDP's intrinsic trait is that the targets are concealed in their surroundings, thus fully perceiving them requires fine-grained representations, prior knowledge, auxiliary reasoning, etc. The contributions of this review are three-fold: (i) We introduce the scope, characteristics, and challenges specific to CDP tasks and emphasize their essential differences from generic vision tasks. (ii) We develop a taxonomy based on concealment counteracting to summarize deep learning efforts in CDP through experiments on three tasks. We compare 25 state-of-the-art methods across 12 widely used concealed datasets. (iii) We discuss the potential applications of CDP in the large model era and summarize 6 potential research directions. We offer perspectives for the future development of CDP by constructing a large-scale multimodal instruction fine-tuning dataset, CvpINST, and a concealed visual perception agent, CvpAgent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。