用合成数据私有化训练专用模型,用户可控隐私与效果平衡
SpinML: Customized Synthetic Data Generation for Private Training of Specialized ML Models
- 服务器基于少量用户参考图生成定制化合成图像用于模型训练
- 在三个任务中提升专用模型性能,且不泄露用户隐私数据
- 支持对象级精细控制,用户可按偏好调节隐私与数据效用
针对智能设备上部署的专用机器学习模型,因缺乏公开标注数据和用户隐私顾虑导致训练困难。本文提出 SpinML 系统,通过仅使用用户提供的少量清洗后的参考图像,由服务器生成定制化合成图像,实现对专用模型的私有化训练。该系统提供细粒度的对象级控制,使用户可根据自身隐私偏好,在生成数据的隐私性与实用性之间灵活权衡。在三个专用模型训练任务上的实验表明,SpinML 可有效提升模型性能,同时保障用户隐私,无需访问原始私有数据。
原文摘要 · Abstract (English)
Specialized machine learning (ML) models tailored to users needs and requests are increasingly being deployed on smart devices with cameras, to provide personalized intelligent services taking advantage of camera data. However, two primary challenges hinder the training of such models: the lack of publicly available labeled data suitable for specialized tasks and the inaccessibility of labeled private data due to concerns about user privacy. To address these challenges, we propose a novel system SpinML, where the server generates customized Synthetic image data to Privately traIN a specialized ML model tailored to the user request, with the usage of only a few sanitized reference images from the user. SpinML offers users fine-grained, object-level control over the reference images, which allows user to trade between the privacy and utility of the generated synthetic data according to their privacy preferences. Through experiments on three specialized model training tasks, we demonstrate that our proposed system can enhance the performance of specialized models without compromising users privacy preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。