arXiv:2605.10229cs.CVcs.CY2026-05中稿 · ICML

构建10万张细粒度隐私图像数据集,提升实时视觉隐私检测能力

VPD-100K: Towards Generalizable and Fine-grained Visual Privacy Protection

论文配图:VPD-100K: Towards Generalizable and Fine-grained Visual Privacy Protection
图 1 · 摘自论文原文
  • 构建四领域细粒度标注的10万张图像数据集
  • 包含超19万实例,覆盖小目标与复杂场景的长尾分布
  • 提出频域注意力模块,更好捕捉隐蔽隐私信息

视觉数据泛滥时代,隐私保护需求日益迫切,对高效稳健的隐私检测算法提出更高要求。然而现有鲁棒检测模型受限于缺乏全面的数据集。当前隐私数据集普遍存在规模有限、标注粗粒度、领域覆盖窄等问题,难以反映真实环境中敏感信息的复杂细节。为此,我们提出大规模细粒度视觉隐私数据集VPD-100K,支持通用隐私检测。该数据集建立涵盖四类核心领域的分类体系:人体存在、屏幕内个人身份信息(PII)、物理标识符和位置指示器,共包含10万张图像,33个细粒度类别,超过19万个体素实例。统计分析表明,数据集具有长尾分布、小物体尺度及高视觉复杂性特征,特别适用于直播等非受限场景下的实时隐私泄露检测。此外,我们设计了一种频域增强型轻量级模块,结合频域注意力融合与自适应谱门机制,突破传统空间像素强度限制,更有效捕捉敏感信息的细微特征。在多样图像与视频流基准上的大量实验一致验证了VPD-100K数据集与优化频域机制的有效性。代码与数据集已公开于https://vpd-100k.github.io/。

原文摘要 · Abstract (English)

Privacy protection has become a critical requirement in the era of ubiquitous visual data sharing, imposing higher demands on efficient and robust privacy detection algorithms. However, current robust detection models are severely hindered by the lack of comprehensive datasets. Existing privacy-oriented datasets often suffer from limited scale, coarse-grained annotations, and narrow domain coverage, failing to capture the intricate details of sensitive information in realworld environments. To bridge this gap, we present a large-scale, fine-grained Visual Privacy Dataset (VPD-100K), designed to facilitate generalized privacy detection. We establish a holistic taxonomy comprising four primary domains: Human Presence, On-Screen Personally Identifiable Information (PII), Physical Identifiers, and Location Indicators, containing 100,000 images annotated with 33 fine-grained classes and over 190,000 object instances. Statistical analysis reveals that our dataset features long-tailed distributions, small object scales, and high visual complexity. These characteristics make the dataset particularly valuable for demanding, unconstrained applications such as live streaming, where actors frequently face unintentional, realtime information leakage. Furthermore, we design an effective frequency-enhanced lightweight module consisting of frequency-domain attention fusion and adaptive spectral gating mechanism that breaks the limitations of spatial pixel intensity to better capture the subtle details of sensitive information. Extensive experiments conducted on both diverse image and streaming videos benchmarks consistently demonstrate the effectiveness of our VPD-100K dataset and the wellcurated frequency mechanism. The code and dataset are available at https://vpd-100k.github.io/.

视觉隐私数据集细粒度检测频域感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。