VLMs在隐私属性识别上与人类标注高度一致,还能发现人类遗漏的细节。
Which private attributes do VLMs agree on and predict well?
- 零样本评估开源视觉语言模型对隐私属性的识别能力。
- VLMs比人类更常预测隐私属性存在,且在高一致性时能补全人工遗漏。
- 适合大规模图像数据集的隐私标注辅助,尤其适用于需快速筛选的场景。
视觉语言模型(VLMs)常用于图像中视觉属性的零样本检测。本文对开源VLMs在隐私相关属性识别上的表现进行零样本评估。我们识别出VLMs具有强一致性的属性,并分析了人类与VLM标注不一致的情况。结果显示,与人类标注对比,VLMs倾向于更频繁地预测隐私属性的存在。此外,在VLMs间达成高一致性的案例中,它们能补充人类标注,发现被忽略的属性。这表明VLMs在大规模图像数据集的隐私标注中具有重要应用潜力。
原文摘要 · Abstract (English)
Visual Language Models (VLMs) are often used for zero-shot detection of visual attributes in the image. We present a zero-shot evaluation of open-source VLMs for privacy-related attribute recognition. We identify the attributes for which VLMs exhibit strong inter-annotator agreement, and discuss the disagreement cases of human and VLM annotations. Our results show that when evaluated against human annotations, VLMs tend to predict the presence of privacy attributes more often than human annotators. In addition to this, we find that in cases of high inter-annotator agreement between VLMs, they can complement human annotation by identifying attributes overlooked by human annotators. This highlights the potential of VLMs to support privacy annotations in large-scale image datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。