用机器学习自动匿名化执法文档,减少人工成本且保留可追溯性。
Anonymization of Documents for Law Enforcement with Machine Learning
- 基于一个参考样本,结合自监督模型实现高效精准去敏。
- 相比纯自动系统和简单复制粘贴方案,去敏更准确、覆盖更全。
- 适合需批量处理敏感文档的执法与司法机构使用。
随着数据驱动方法在执法等敏感信息处理领域应用日益广泛,机构必须持续投入以满足数据保护要求。本文提出一种自动匿名化扫描文档图像的系统,在降低人工成本的同时保障合规性。该方法通过最小化自动遮蔽区域,兼顾去敏后仍可进行后续法证分析的需求:利用自监督图像模型对参考文档进行实例检索,仅需一个手动匿名化的样本即可高效处理同类型全部文档,显著缩短处理时间。我们在一个手工标注的真实去敏数据集上验证,该方法优于纯自动去敏系统及简单的参考样本复制粘贴方案。
原文摘要 · Abstract (English)
The steadily increasing utilization of data-driven methods and approaches in areas that handle sensitive personal information such as in law enforcement mandates an ever increasing effort in these institutions to comply with data protection guidelines. In this work, we present a system for automatically anonymizing images of scanned documents, reducing manual effort while ensuring data protection compliance. Our method considers the viability of further forensic processing after anonymization by minimizing automatically redacted areas by combining automatic detection of sensitive regions with knowledge from a manually anonymized reference document. Using a self-supervised image model for instance retrieval of the reference document, our approach requires only one anonymized example to efficiently redact all documents of the same type, significantly reducing processing time. We show that our approach outperforms both a purely automatic redaction system and also a naive copy-paste scheme of the reference anonymization to other documents on a hand-crafted dataset of ground truth redactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。