arXiv:2506.23663cs.CVcs.LG2025-06被引 2

提出新框架Deepbench,评估视觉语言模型在真实场景下的鲁棒性。

On the Domain Robustness of Contrastive Vision-Language Models

  • 用大语言模型生成特定场景的图像噪声,无需标注数据
  • 在6个真实场景中测试,发现模型鲁棒性差异显著
  • 适合关注模型落地可靠性的研究人员

在真实视觉语言应用中,从业者越来越多依赖大规模预训练基础模型,而非定制解决方案,尽管其训练数据和过程透明度有限。尽管这些模型在通用基准上表现优异,但在特定领域迁移时性能可能显著下降,如特殊成像条件或环境变化。本文提出Deepbench框架,用于评估视觉语言模型(VLMs)在特定领域的鲁棒性。该框架利用大语言模型(LLM)生成与具体部署场景相关的、具上下文意识的图像退化,无需标注数据。我们在六个真实世界领域对多种对比式视觉语言架构及其变体进行了评估,观察到鲁棒性存在显著差异,凸显了针对性、领域感知评估的必要性。Deepbench已开源,以支持领域感知鲁棒性评估的进一步研究。

原文摘要 · Abstract (English)

In real-world vision-language applications, practitioners increasingly rely on large, pretrained foundation models rather than custom-built solutions, despite limited transparency regarding their training data and processes. While these models achieve impressive performance on general benchmarks, their effectiveness can decline notably under specialized domain shifts, such as unique imaging conditions or environmental variations. In this work, we introduce Deepbench, a framework designed to assess domain-specific robustness of vision-language models (VLMs). Deepbench leverages a large language model (LLM) to generate realistic, context-aware image corruptions tailored to specific deployment domains without requiring labeled data. We evaluate a range of contrastive vision-language architectures and architectural variants across six real-world domains and observe substantial variability in robustness, highlighting the need for targeted, domain-aware evaluation. Deepbench is released as open-source software to support further research into domain-aware robustness assessment.

视觉语言模型鲁棒性评估领域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。