arXiv:2412.00067cs.CVcs.LG2024-12

基于场景图的物体级数据删除,精准移除图像中特定对象而不影响其他内容。

Targeted Therapy in Data Removal: Object Unlearning Based on Scene Graphs

  • 用场景图表达语义关系,将删除请求转为可执行操作。
  • 在图像重建与合成任务中保持未删除对象质量,删除效果更精准。
  • 适合需要细粒度隐私保护的图像生成服务,如含人脸的图片去标识化。

用户可能无意中将个人身份信息(PII)上传至机器学习即服务(MLaaS)平台。根据GDPR和COPPA等法规,用户有权要求删除其数据。因此,服务方需高效实现特定数据点的影响清除。传统方法仅支持整样本或全特征的删除(样本/特征遗忘),难以应对样本内特定物体的精细化删除需求。为此,本文提出一种基于场景图的物体级遗忘框架。该框架利用富含语义的场景图,将遗忘请求转化为具体操作,确保生成图像的整体语义完整性,仅移除指定物体。同时采用影响函数缓解高计算开销。在图像重建与合成任务中验证,所提方法在保留未请求样本质量方面优于传统样本与特征遗忘方法。本工作通过提升遗忘粒度,实现在不牺牲数据集整体可用性的前提下,精准删除特定对象信息,解决关键隐私问题。

原文摘要 · Abstract (English)

Users may inadvertently upload personally identifiable information (PII) to Machine Learning as a Service (MLaaS) providers. When users no longer want their PII on these services, regulations like GDPR and COPPA mandate a right to forget for these users. As such, these services seek efficient methods to remove the influence of specific data points. Thus the introduction of machine unlearning. Traditionally, unlearning is performed with the removal of entire data samples (sample unlearning) or whole features across the dataset (feature unlearning). However, these approaches are not equipped to handle the more granular and challenging task of unlearning specific objects within a sample. To address this gap, we propose a scene graph-based object unlearning framework. This framework utilizes scene graphs, rich in semantic representation, transparently translate unlearning requests into actionable steps. The result, is the preservation of the overall semantic integrity of the generated image, bar the unlearned object. Further, we manage high computational overheads with influence functions to approximate the unlearning process. For validation, we evaluate the unlearned object's fidelity in outputs under the tasks of image reconstruction and image synthesis. Our proposed framework demonstrates improved object unlearning outcomes, with the preservation of unrequested samples in contrast to sample and feature learning methods. This work addresses critical privacy issues by increasing the granularity of targeted machine unlearning through forgetting specific object-level details without sacrificing the utility of the whole data sample or dataset feature.

数据删除图像生成隐私保护场景图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。