通过随机替换数据保护隐私,防止信息泄露。
PASS: Private Attributes Protection with Stochastic Data Substitution
- 用概率替换样本,避免直接删除敏感信息
- 在多模态数据上验证,隐私保护效果显著
- 适用于图像、语音、传感信号等场景
随着机器学习服务对用户数据需求的增长,数据中可能包含与服务无关的个人隐私信息。现有方法多通过移除敏感属性来保护隐私,但本文理论与实证表明,基于对抗训练的策略存在严重漏洞。为此,我们提出PASS方法:基于概率随机替换原始样本,并采用源于信息论目标的新型损失函数,实现保留数据效用的同时保护隐私。在面部图像、人体活动传感信号和语音记录等多种模态数据集上的全面评估表明,PASS在隐私保护与任务性能间取得良好平衡,具备强泛化能力。
原文摘要 · Abstract (English)
The growing Machine Learning (ML) services require extensive collections of user data, which may inadvertently include people's private information irrelevant to the services. Various studies have been proposed to protect private attributes by removing them from the data while maintaining the utilities of the data for downstream tasks. Nevertheless, as we theoretically and empirically show in the paper, these methods reveal severe vulnerability because of a common weakness rooted in their adversarial training based strategies. To overcome this limitation, we propose a novel approach, PASS, designed to stochastically substitute the original sample with another one according to certain probabilities, which is trained with a novel loss function soundly derived from information-theoretic objective defined for utility-preserving private attributes protection. The comprehensive evaluation of PASS on various datasets of different modalities, including facial images, human activity sensory signals, and voice recording datasets, substantiates PASS's effectiveness and generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。