arXiv:2505.04835cs.CV2025-05中稿 · CVPR被引 5

对比真实与合成退化数据,验证合成数据能否可靠评估模型鲁棒性。

Are Synthetic Corruptions A Reliable Proxy For Real-World Corruptions?

  • 构建最大规模语义分割模型鲁棒性评测,对比真实与合成退化数据
  • 发现合成与真实退化下平均性能高度相关(相关系数达0.92)
  • 揭示特定退化类型中合成数据的适用边界,指导测试设计

深度学习模型在实际应用中仍易受分布偏移影响,尤其因天气和光照变化。收集多样化的现实数据以测试模型鲁棒性成本高昂,因此合成退化成为替代方案。但合成退化是否能可靠代表真实退化?我们开展了最大规模的语义分割模型评测,比较模型在真实退化数据集与合成退化数据集上的表现。结果表明,两类数据下的平均性能具有强相关性(相关系数0.92),支持使用合成退化进行鲁棒性评估。进一步分析特定退化类型的关联性,揭示了合成退化在哪些场景下能有效模拟真实退化。开源代码:https://github.com/shashankskagnihotri/benchmarking_robustness/tree/segmentation_david/semantic_segmentation

原文摘要 · Abstract (English)

Deep learning (DL) models are widely used in real-world applications but remain vulnerable to distribution shifts, especially due to weather and lighting changes. Collecting diverse real-world data for testing the robustness of DL models is resource-intensive, making synthetic corruptions an attractive alternative for robustness testing. However, are synthetic corruptions a reliable proxy for real-world corruptions? To answer this, we conduct the largest benchmarking study on semantic segmentation models, comparing performance on real-world corruptions and synthetic corruptions datasets. Our results reveal a strong correlation in mean performance, supporting the use of synthetic corruptions for robustness evaluation. We further analyze corruption-specific correlations, providing key insights to understand when synthetic corruptions succeed in representing real-world corruptions. Open-source Code: https://github.com/shashankskagnihotri/benchmarking_robustness/tree/segmentation_david/semantic_segmentation

模型鲁棒性合成数据语义分割退化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。