arXiv:2505.06621cs.LGcs.CV2025-05被引 6

用代理任务训练儿童色情内容检测模型,避免接触真实敏感数据

Minimizing Risk Through Minimizing Model-Data Interaction: A Protocol For Relying on Proxy Tasks When Designing Child Sexual Abuse Imagery Detection Models

  • 用代理任务替代真实数据训练模型,减少模型与敏感数据交互
  • 首次实现少样本室内场景分类,在真实数据集上表现良好
  • 提出可复用的协议,适合安全合规的AI检测系统设计者

儿童性虐待影像(CSAI)的传播是当今世界日益严峻的问题,受害者二次受创,海量非法影像使执法机构面临繁重的手动分类负担。为减轻压力,研究者尝试自动化数据筛选与检测,但数据敏感性导致真实数据难以获取,模型训练必须避免与真实数据直接交互,以防泄露。本文提出“代理任务”概念,指在不使用真实CSAI数据的情况下,用于训练检测模型的替代任务。基于此,系统回顾现有文献,并提出一套结合执法机构持续输入的协议,指导更有效的自动化系统设计。最后,将该协议应用于首次开展的少样本室内场景分类任务,在真实数据集上构建出性能优异的模型,且其权重从未在敏感数据上训练。

原文摘要 · Abstract (English)

The distribution of child sexual abuse imagery (CSAI) is an ever-growing concern of our modern world; children who suffered from this heinous crime are revictimized, and the growing amount of illegal imagery distributed overwhelms law enforcement agents (LEAs) with the manual labor of categorization. To ease this burden researchers have explored methods for automating data triage and detection of CSAI, but the sensitive nature of the data imposes restricted access and minimal interaction between real data and learning algorithms, avoiding leaks at all costs. In observing how these restrictions have shaped the literature we formalize a definition of "Proxy Tasks", i.e., the substitute tasks used for training models for CSAI without making use of CSA data. Under this new terminology we review current literature and present a protocol for making conscious use of Proxy Tasks together with consistent input from LEAs to design better automation in this field. Finally, we apply this protocol to study -- for the first time -- the task of Few-shot Indoor Scene Classification on CSAI, showing a final model that achieves promising results on a real-world CSAI dataset whilst having no weights actually trained on sensitive data.

内容检测代理任务少样本学习隐私安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。