arXiv:2603.17357cs.CRcs.AI2026-03被引 2

构建首个网页截图个人隐私识别基准,提升敏感信息检测精度。

WebPII: Benchmarking Visual PII Detection for Computer-Use Agents

  • 基于视觉大模型生成4.5万张带标注电商界面图,支持细粒度隐私识别。
  • 在部分填写表单场景下实现提前识别,检测准确率超基线一倍以上。
  • 适用于开发隐私保护型自动化工具,适合关注数据安全的研究者。

计算机使用代理带来新隐私风险:训练数据来自真实网站必然包含敏感信息,云端推理则暴露用户截图。检测网页截图中的个人可识别信息对隐私保护部署至关重要,但目前尚无公开基准。我们提出WebPII,一个包含44,865张标注的电商界面合成数据集,具备三大特性:扩展的隐私信息分类体系(含可重识别的交易级标识符)、支持部分填写表单的前瞻检测能力,以及基于视觉大模型的可扩展界面生成方法。实验验证这些设计提升了跨多样化界面的布局无关检测性能,并增强对未见页面类型的泛化能力。我们训练了WebRedact模型,实现在20ms实时CPU延迟下,文本提取准确率(mAP@50)达0.753,较基线0.357提升超一倍。数据集与模型已开源,助力隐私保护型计算机使用研究。

原文摘要 · Abstract (English)

Computer use agents create new privacy risks: training data collected from real websites inevitably contains sensitive information, and cloud-hosted inference exposes user screenshots. Detecting personally identifiable information in web screenshots is critical for privacy-preserving deployment, but no public benchmark exists for this task. We introduce WebPII, a fine-grained synthetic benchmark of 44,865 annotated e-commerce UI images designed with three key properties: extended PII taxonomy including transaction-level identifiers that enable reidentification, anticipatory detection for partially-filled forms where users are actively entering data, and scalable generation through VLM-based UI reproduction. Experiments validate that these design choices improve layout-invariant detection across diverse interfaces and generalization to held-out page types. We train WebRedact to demonstrate practical utility, more than doubling text-extraction baseline accuracy (0.753 vs 0.357 mAP@50) at real-time CPU latency (20ms). We release the dataset and model to support privacy-preserving computer use research.

隐私检测视觉大模型数据安全基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。