统合视觉检测新范式,推动机器感知从封闭到开放世界演进
Towards Open World Detection: A Survey
- 提出开放世界检测框架,融合无类别目标检测与通用感知模型
- 梳理从早期显著性检测到大语言模型的演进脉络与关键技术
- 适合关注下一代视觉系统通用性与可扩展性的研究者
几十年来,计算机视觉致力于使机器能够感知外部世界。早期局限催生了高度专业化的细分领域。随着各任务取得进展,更复杂的感知任务逐渐涌现。本文梳理这些任务的融合发展路径,提出开放世界检测(OWD)这一统称,用于整合视觉领域中无类别且通用的目标检测模型。从基础视觉子领域的历史出发,涵盖当前前沿的核心概念、方法与数据集。内容涉及早期显著性检测、前景/背景分离、分布外检测,直至开放世界目标检测、零样本检测和视觉大语言模型(VLLMs)。探讨这些子领域间的重叠与融合趋势,展望其未来可能统一为单一感知范式。
原文摘要 · Abstract (English)
For decades, Computer Vision has aimed at enabling machines to perceive the external world. Initial limitations led to the development of highly specialized niches. As success in each task accrued and research progressed, increasingly complex perception tasks emerged. This survey charts the convergence of these tasks and, in doing so, introduces Open World Detection (OWD), an umbrella term we propose to unify class-agnostic and generally applicable detection models in the vision domain. We start from the history of foundational vision subdomains and cover key concepts, methodologies and datasets making up today's state-of-the-art landscape. This traverses topics starting from early saliency detection, foreground/background separation, out of distribution detection and leading up to open world object detection, zero-shot detection and Vision Large Language Models (VLLMs). We explore the overlap between these subdomains, their increasing convergence, and their potential to unify into a singular domain in the future, perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。