提出新任务与数据集,评估视觉区域目标检测能力
RegionDet: A Benchmark for Region Detection Beyond Object Instances

- 扩展目标检测至区域级任务,涵盖8类非个体视觉区域
- 零样本模型表现差,暴露现有模型的物体中心偏见
- 适合关注场景理解、视觉语言模型泛化能力的研究者
传统目标检测聚焦离散、边界清晰的物体实例,但现实场景中许多视觉目标是基于视觉状态、场景上下文、物体关系和人类活动定义的区域,如施工区、破损路面、排队、群体对话、摊贩区等。现有检测基准多围绕物体实例构建,难以系统评估此类区域目标。为此,我们提出区域检测任务,并构建了包含8类区域(施工、过街、损坏、排队、交谈、摊贩、等待、行走)的RegionDet基准,采用COCO风格的边界框标注与评估协议。我们系统评估了代表性闭集与零样本/开放词汇检测器,结果表明闭集模型在监督下可部分学习区域模式,而零样本模型表现严重受限,揭示当前视觉-语言模型存在强烈的物体中心偏见。进一步分析指出关键挑战:边界线索弱、上下文依赖强、关系层级理解不足。RegionDet将公开发布。
原文摘要 · Abstract (English)
Object detection is a fundamental task in computer vision and has achieved remarkable progress on standard benchmarks by localizing discrete and well-bounded object instances. However, many visual targets in real-world scenarios are not individual objects, but regions defined by visual states, scene context, object relations, and human activities, such as construction areas, damaged road regions, queues, group conversations, and vendor regions. Existing detection benchmarks are mainly built around object instances, providing limited support for systematically evaluating such region targets. To address this gap, we introduce Region Detection, a task that extends conventional object detection beyond object instances, and construct RegionDet, a benchmark for region target localization. RegionDet contains eight region categories, including Construction, Crossing, Damage, Queuing, Talking, Vendor, Waiting, and Walking, with COCO-style bounding-box annotations and evaluation protocols. We systematically evaluate representative closed-set and zero-shot/open-vocabulary detectors on RegionDet. Results show that closed-set detectors can partially learn region-level patterns under supervision, while zero-shot/open-vocabulary detectors struggle severely, revealing the strong object-centric bias of current vision-language detectors. Further analyses highlight key challenges in Region Detection, including weak boundary cues, strong context dependency, and insufficient relation-level region understanding. The RegionDet will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。