arXiv:2608.06841cs.CV2026-08

拓展目标检测范围,让模型识别天空、道路等非实体视觉元素。

ECAD: Expanding Class-Agnostic Detection Beyond Thing-Centric Objectness

论文配图:ECAD: Expanding Class-Agnostic Detection Beyond Thing-Centric Objectness
图 1 · 摘自论文原文
  • 提出新检测范式ECAD,识别传统方法忽略的非实体视觉元素。
  • 在BTCO-Bench上超越主流检测器,在跨域场景下提升显著。
  • 适合关注场景理解与空间推理的研究者使用。

目标检测是视觉感知的基础任务,提供结构化区域表示以支持识别、定位、推理与交互。然而,现有检测范式主要基于‘实体中心’的对象观念,仅训练模型定位离散且可数的对象实例。因此,诸如天空、道路、草地、水域和运动场等语义重要的视觉元素常被归入背景,影响场景理解与空间推理。本文提出扩展的类别无关检测(ECAD),旨在发现超越传统实体对象的类别无关视觉候选。为此,构建了涵盖真实世界与跨域场景的BTCO-Bench基准,包含类别无关框标注。进一步提出基于冻结DINOv3编码器的轻量级DETR检测器ECADet,引入几何感知专家回归(GAER)与原型引导查询调制(PGQM),分别提升多样视觉元素的定位精度与对象性估计能力。大量实验表明,ECADet在BTCO-Bench上持续优于代表性类别无关与提案类检测器,验证了扩展对象性发现的有效性。代码与基准将公开。

原文摘要 · Abstract (English)

Object detection is a fundamental task in visual perception, providing structured region representations for recognition, grounding, reasoning, and interaction. However, existing detection paradigms largely inherit a thing-centric notion of objectness, where detectors are mainly trained to localize discrete and countable object instances. Consequently, many semantically meaningful visual elements, such as sky, road, grassland, water, and sports courts, are often absorbed into the background despite their importance for scene understanding and spatial reasoning. In this paper, we formulate Expanded Class-Agnostic Detection (ECAD), a new setting that aims to discover category-agnostic visual candidates beyond conventional thing-centric objects. To support this setting, we construct BTCO-Bench, a Beyond Thing-Centric Objectness benchmark with category-agnostic box annotations covering both real-world and cross-domain scenarios. We further propose ECADet, a lightweight DETR-based detector built upon a frozen DINOv3 encoder, and introduce Geometry-Aware Expert Regression (GAER) and Prototype-Guided Query Modulation (PGQM) to improve localization and objectness estimation for diverse visual elements, respectively. Extensive experiments show that ECADet consistently outperforms representative class-agnostic and proposal-based detectors on BTCO-Bench, demonstrating the effectiveness of expanded objectness discovery. Code and benchmark will be released.

目标检测场景理解扩散模型图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。