大模型在零样本隐私分类中抗干扰强但速度慢,不如专用小模型精准。
On the Robustness of Vision-Language Models in Zero-shot Privacy Classification
- 用三个开源大视觉语言模型做零样本隐私分类,对比专用小模型表现。
- 大模型对压缩、噪声等干扰仍保持较高鲁棒性,但准确率低20%以上。
- 适合追求泛化能力的场景,不适用于实时隐私处理需求。
文档理解系统需要能准确识别敏感视觉内容的多模态模型,即使面对图像退化也需可靠。指令跟随型大型视觉-语言模型(VLMs)预期可在无特定适配的情况下跨领域、跨任务泛化。本文系统分析了VLMs在零样本隐私分类中的可靠性。我们在两个公开标准基准上评估并比较了三种开源VLMs与专为隐私设计的模型的分类性能。评估涵盖压缩、光照变化、随机噪声等扰动下的鲁棒性,并分析推理速度与参数量,以支持在隐私敏感文档处理流程中的部署。结果表明,大型VLMs对输入扰动具有较强鲁棒性,但准确率低于专用小模型(差值>20%),且推理速度慢得多。单纯扩大规模不足以提升效果,凸显了专为隐私分类设计模型的优势。
原文摘要 · Abstract (English)
Automatic systems for document understanding require multimodal models that accurately identify sensitive visual content, even in the presence of image degradations. Instruction-following large Vision-Language Models (VLMs) are expected to generalise across domains and tasks without requiring any specific adaptation. In this work, we systematically analyse whether VLMs can be used reliably for image privacy classification in a zero-shot setup. We evaluate and compare the classification performance of three open-source VLMs against purposely built models on two public, standard benchmarks. We assess robustness to image degradations caused by perturbations such as compression, light variations, and random noise, and analyse inference speed and parameter count to deploy VLMs in privacy-aware document processing pipelines. Our results show that large VLMs are robust to input perturbations but are less accurate (and much slower) than smaller privacy models. Scaling alone is not sufficient, highlighting the advantages of models specifically designed for privacy classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。