arXiv:2510.05581cs.LGcs.CR2025-10

提出可保护隐私的表格数据嵌入共享机制,支持任意模型在服务端使用。

Power Mechanism: Private Tabular Representation Release for Model Agnostic Consumption

  • 联合设计隐私编码与效用生成网络,实现嵌入的严格差分隐私
  • 仅需一轮通信,客户端计算量低于传统方法
  • 嵌入结果对服务端模型类型无感,兼容深度学习/树模型

传统协同学习依赖客户端与服务器间共享模型权重,但基于数据生成的嵌入(activations)共享更具资源效率。现有方法已支持权重的差分隐私保护,但针对嵌入的隐私机制尚缺。本文提出一种联合优化框架:训练一个隐私编码网络和小型效用生成网络,使最终生成的嵌入具备正式差分隐私保证。这些经隐私处理的嵌入被发送至更强的服务器,由其执行后处理以提升任务准确率。实验表明,该协同与隐私学习的共设计方案仅需一轮隐私通信,且客户端计算开销更低。所共享的嵌入对服务端使用的模型类型(如深度学习、随机森林或XGBoost)完全无关,可广泛用于各类下游任务。

原文摘要 · Abstract (English)

Traditional collaborative learning approaches are based on sharing of model weights between clients and a server. However, there are advantages to resource efficiency through schemes based on sharing of embeddings (activations) created from the data. Several differentially private methods were developed for sharing of weights while such mechanisms do not exist so far for sharing of embeddings. We propose Ours to learn a privacy encoding network in conjunction with a small utility generation network such that the final embeddings generated from it are equipped with formal differential privacy guarantees. These privatized embeddings are then shared with a more powerful server, that learns a post-processing that results in a higher accuracy for machine learning tasks. We show that our co-design of collaborative and private learning results in requiring only one round of privatized communication and lesser compute on the client than traditional methods. The privatized embeddings that we share from the client are agnostic to the type of model (deep learning, random forests or XGBoost) used on the server in order to process these activations to complete a task.

隐私计算嵌入共享差分隐私模型无关

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。