arXiv:2505.18015cs.CVcs.LG2025-05被引 2

构建分割与检测模型的可靠性评估基准,揭示当前顶尖模型的系统性弱点。

SemSegBench & DetecBench: Benchmarking Reliability and Generalization Beyond Classification

  • 提出针对分割和检测任务的可靠性评测工具集
  • 在4个数据集上测试76个分割模型,2个数据集上测试61个检测器
  • 公开全部6139次评估结果,助力安全关键场景模型研究

深度学习的可靠性与泛化能力主要在图像分类领域被研究,但实际安全关键应用涉及更广泛的语义任务,如语义分割和目标检测,这些任务使用多样化的专用模型架构。为推动分割与检测模型的鲁棒设计,本文提出SEMSEGBENCH和DETECBENCH两个基准评测工具,对当前最全面的可靠性与泛化性能评估。我们评估了76个分割模型(跨4个数据集)和61个目标检测器(跨2个数据集),在多种对抗攻击和常见损坏下的表现。结果揭示了当前顶级模型存在系统性弱点,并发现架构、主干网络与模型容量对性能的影响趋势。所有评测结果(共6139次)均已开源至GitHub(https://github.com/shashankskagnihotri/benchmarking_reliability_generalization),期望推动超越分类任务的模型可靠性研究。

原文摘要 · Abstract (English)

Reliability and generalization in deep learning are predominantly studied in the context of image classification. Yet, real-world applications in safety-critical domains involve a broader set of semantic tasks, such as semantic segmentation and object detection, which come with a diverse set of dedicated model architectures. To facilitate research towards robust model design in segmentation and detection, our primary objective is to provide benchmarking tools regarding robustness to distribution shifts and adversarial manipulations. We propose the benchmarking tools SEMSEGBENCH and DETECBENCH, along with the most extensive evaluation to date on the reliability and generalization of semantic segmentation and object detection models. In particular, we benchmark 76 segmentation models across four datasets and 61 object detectors across two datasets, evaluating their performance under diverse adversarial attacks and common corruptions. Our findings reveal systematic weaknesses in state-of-the-art models and uncover key trends based on architecture, backbone, and model capacity. SEMSEGBENCH and DETECBENCH are open-sourced in our GitHub repository (https://github.com/shashankskagnihotri/benchmarking_reliability_generalization) along with our complete set of total 6139 evaluations. We anticipate the collected data to foster and encourage future research towards improved model reliability beyond classification.

模型可靠性分割检测基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。