首次系统评估视觉大模型对常见图像编辑的鲁棒性,发现普遍不稳健。
Robustness of Vision Foundation Models to Common Perturbations
- 提出三种鲁棒性度量方法并验证其数学性质
- 六款工业级模型在九类扰动下性能普遍下降
- 可预测下游任务受损程度,支持针对性优化
视觉基础模型将图像映射为嵌入向量,而常见编辑操作(如JPEG压缩、亮度/对比度调整)会改变嵌入,影响下游任务表现。本文首次系统研究基础模型对这类扰动的鲁棒性,提出三种鲁棒性度量并分析其五项数学性质。基于这些度量,评估了六款工业级模型(OpenAI、Meta)在九类常见扰动下的表现,发现模型普遍缺乏鲁棒性。同时表明,常见扰动会降低下游应用性能(如分类准确率),且鲁棒性指标可有效预测性能损失。最后提出一种微调方法,在不牺牲功能的前提下提升模型鲁棒性。
原文摘要 · Abstract (English)
A vision foundation model outputs an embedding vector for an image, which can be affected by common editing operations (e.g., JPEG compression, brightness, contrast adjustments). These common perturbations alter embedding vectors and may impact the performance of downstream tasks using these embeddings. In this work, we present the first systematic study on foundation models' robustness to such perturbations. We propose three robustness metrics and formulate five desired mathematical properties for these metrics, analyzing which properties they satisfy or violate. Using these metrics, we evaluate six industry-scale foundation models (OpenAI, Meta) across nine common perturbation categories, finding them generally non-robust. We also show that common perturbations degrade downstream application performance (e.g., classification accuracy) and that robustness values can predict performance impacts. Finally, we propose a fine-tuning approach to improve robustness without sacrificing utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。