通用搜索模型意外学会识别人脸,需警惕执法中的隐私风险
Emergent AI Surveillance: Overlearned Person Re-Identification and Its Mitigation in Law Enforcement Context
- 模型在无真人数据下仍能识别特定个体,源于过度学习
- 结合去标识与混淆损失可将识别人准确率压至2%以下
- 适合关注AI监管与数据安全的政策制定者与工程师
通用实例搜索模型可显著降低刑事案件中分析海量监控视频的人工成本,通过检索目标对象。然而我们的研究发现,这些模型在未使用人类主体数据的情况下,仍会因过度学习而具备识别特定个体的能力,引发基于个人数据的识别与画像风险。目前尚无明确标准实现有效去标识化。我们评估了两种技术防护措施:索引排除与混淆损失。实验表明,联合使用可使人物重识别准确率降至2%以下,同时保持非人物对象检索性能达82%。但发现关键漏洞,如利用部分人体图像仍可绕过防护。该研究揭示了人工智能治理与数据保护交叉领域的紧迫问题:如何分类和监管具有隐性识别能力的系统?在看似无害的应用中,应建立何种技术标准以防止识别能力的涌现?
原文摘要 · Abstract (English)
Generic instance search models can dramatically reduce the manual effort required to analyze vast surveillance footage during criminal investigations by retrieving specific objects of interest to law enforcement. However, our research reveals an unintended emergent capability: through overlearning, these models can single out specific individuals even when trained on datasets without human subjects. This capability raises concerns regarding identification and profiling of individuals based on their personal data, while there is currently no clear standard on how de-identification can be achieved. We evaluate two technical safeguards to curtail a model's person re-identification capacity: index exclusion and confusion loss. Our experiments demonstrate that combining these approaches can reduce person re-identification accuracy to below 2% while maintaining 82% of retrieval performance for non-person objects. However, we identify critical vulnerabilities in these mitigations, including potential circumvention using partial person images. These findings highlight urgent regulatory questions at the intersection of AI governance and data protection: How should we classify and regulate systems with emergent identification capabilities? And what technical standards should be required to prevent identification capabilities from developing in seemingly benign applications?
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。