首个评估视觉语言模型在不同可见度下隐私泄露的基准,发现高可见性用户更易被识别。
PII-VisBench: Evaluating Personally Identifiable Information Safety in Vision Language Models Along a Continuum of Visibility
- 构建4000个探针,按在线可见度分四类,评估模型在不同隐私暴露水平下的安全表现
- 高可见度下PII披露率9.10%,低可见度降至5.34%,拒绝率随可见度下降而上升
- 揭示模型家族差异与提示攻击漏洞,推动更具可见度感知的安全设计
视觉语言模型(VLM)正越来越多应用于隐私敏感领域,但现有对个人身份信息(PII)泄露的评估大多将隐私视为静态提取任务,忽视了主体在线存在感——即其数据在网络上的可获取量——对隐私对齐的影响。本文提出PII-VisBench,一个包含4000个独特探针的新基准,用于评估模型在在线可见度连续谱上的安全性。该基准将200名主体按其在线信息的范围和性质划分为高、中、低、零四类可见度。我们基于两个关键指标评估18个开源VLM(规模0.3B–32B):PII探针请求的拒绝率(拒绝率)和非拒绝响应中被标记为含PII的比例(条件性PII披露率)。结果显示,随着主体可见度降低,拒绝率上升,PII披露率下降(从高可见度的9.10%降至低可见度的5.34%)。我们发现模型更倾向于对高可见性主体披露PII,且存在显著的模型族差异与PII类型差异。此外,改写提示和越狱式攻击暴露了模型依赖性的安全缺陷,凸显了开展可见度感知的安全评估与训练干预的必要性。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) are increasingly integrated into privacy-critical domains, yet existing evaluations of personally identifiable information (PII) leakage largely treat privacy as a static extraction task and ignore how a subject's online presence--the volume of their data available online--influences privacy alignment. We introduce PII-VisBench, a novel benchmark containing 4000 unique probes designed to evaluate VLM safety through the continuum of online presence. The benchmark stratifies 200 subjects into four visibility categories: high, medium, low, and zero--based on the extent and nature of their information available online. We evaluate 18 open-source VLMs (0.3B-32B) based on two key metrics: percentage of PII probing queries refused (Refusal Rate) and the fraction of non-refusal responses flagged for containing PII (Conditional PII Disclosure Rate). Across models, we observe a consistent pattern: refusals increase and PII disclosures decrease (9.10% high to 5.34% low) as subject visibility drops. We identify that models are more likely to disclose PII for high-visibility subjects, alongside substantial model-family heterogeneity and PII-type disparities. Finally, paraphrasing and jailbreak-style prompts expose attack and model-dependent failures, motivating visibility-aware safety evaluation and training interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。