构建大规模分布外数据集,评估目标检测与多模态大模型的泛化能力。
COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
- 设计14类自然分布偏移,提供超222万样本与119万标注框
- 发现大模型在分布外任务中性能显著下降,如GPT-4o仅56.7%准确率
- 为检测器和多模态模型提供可复现的分布外评估基准
当前目标检测器在真实场景中遭遇分布偏移时性能严重下降,其分布外(OOD)泛化能力日益受到关注。然而,仍缺乏大规模、细粒度标注的综合性数据集与评测基准来评估复杂任务如目标检测与视觉定位中的分布外泛化。为此,我们提出COUNTS,一个包含对象级标注的大规模分布外数据集。COUNTS涵盖14种自然分布偏移,包含超过222,000个样本和1,196,000个标注边界框。基于COUNTS,我们引入两个新基准:O(OD)2用于评估目标检测器的分布外泛化能力,通过控制训练与测试数据间的分布偏移;OODG则用于评估多模态大语言模型(MLLMs)的视觉定位分布外泛化能力。研究发现,尽管大规模模型和海量预训练数据能显著提升分布内(IID)表现,但在分布外场景下,目标检测器与MLLMs仍存在明显局限与改进空间。在视觉定位任务中,即使先进模型GPT-4o和Gemini-1.5也分别仅达到56.7%和28.0%的准确率。我们希望COUNTS能推动鲁棒目标检测器与多模态大模型在分布偏移下的持续发展与评估。
原文摘要 · Abstract (English)
Current object detectors often suffer significant perfor-mance degradation in real-world applications when encountering distributional shifts. Consequently, the out-of-distribution (OOD) generalization capability of object detectors has garnered increasing attention from researchers. Despite this growing interest, there remains a lack of a large-scale, comprehensive dataset and evaluation benchmark with fine-grained annotations tailored to assess the OOD generalization on more intricate tasks like object detection and grounding. To address this gap, we introduce COUNTS, a large-scale OOD dataset with object-level annotations. COUNTS encompasses 14 natural distributional shifts, over 222K samples, and more than 1,196K labeled bounding boxes. Leveraging COUNTS, we introduce two novel benchmarks: O(OD)2 and OODG. O(OD)2 is designed to comprehensively evaluate the OOD generalization capabilities of object detectors by utilizing controlled distribution shifts between training and testing data. OODG, on the other hand, aims to assess the OOD generalization of grounding abilities in multimodal large language models (MLLMs). Our findings reveal that, while large models and extensive pre-training data substantially en hance performance in in-distribution (IID) scenarios, significant limitations and opportunities for improvement persist in OOD contexts for both object detectors and MLLMs. In visual grounding tasks, even the advanced GPT-4o and Gemini-1.5 only achieve 56.7% and 28.0% accuracy, respectively. We hope COUNTS facilitates advancements in the development and assessment of robust object detectors and MLLMs capable of maintaining high performance under distributional shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。