提出黑盒评估大模型推理效率鲁棒性的新方法,发现微小扰动可使计算开销飙升128倍
VLMInferSlow: Evaluating the Efficiency Robustness of Large Vision-Language Models as a Service
- 在无法访问模型参数的黑盒场景下,通过零阶优化搜索对抗样本
- 微小视觉扰动使VLM推理耗时最高增加128.47%
- 为云服务部署的视觉语言模型提供真实高效的评估工具
视觉语言模型(VLMs)在实际应用中展现出巨大潜力。尽管现有研究主要关注其准确性提升,但效率问题仍被忽视。考虑到许多应用对实时性要求高且VLM推理开销大,效率鲁棒性至关重要。然而,以往研究多基于不切实际的假设——需访问模型架构与参数,这在机器学习即服务(ML-as-a-service)场景中不可行,因VLM通常通过推理API部署。为此,我们提出VLMInferSlow,一种在真实黑盒环境下评估VLM效率鲁棒性的新方法。该方法结合针对VLM推理的细粒度效率建模,并利用零阶优化搜索对抗样本。实验表明,VLMInferSlow生成的对抗图像仅含难以察觉的扰动,却可使计算成本最高提升128.47%。我们希望此研究能引起学术界对VLM效率鲁棒性的重视。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have demonstrated great potential in real-world applications. While existing research primarily focuses on improving their accuracy, the efficiency remains underexplored. Given the real-time demands of many applications and the high inference overhead of VLMs, efficiency robustness is a critical issue. However, previous studies evaluate efficiency robustness under unrealistic assumptions, requiring access to the model architecture and parameters -- an impractical scenario in ML-as-a-service settings, where VLMs are deployed via inference APIs. To address this gap, we propose VLMInferSlow, a novel approach for evaluating VLM efficiency robustness in a realistic black-box setting. VLMInferSlow incorporates fine-grained efficiency modeling tailored to VLM inference and leverages zero-order optimization to search for adversarial examples. Experimental results show that VLMInferSlow generates adversarial images with imperceptible perturbations, increasing the computational cost by up to 128.47%. We hope this research raises the community's awareness about the efficiency robustness of VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。