arXiv:2506.16460cs.LGcs.CR2025-06

黑盒攻击揭示多任务学习共享表示中的隐私泄露风险

Black-Box Privacy Attacks on Shared Representations in Multitask Learning

  • 提出新型黑盒攻击,仅用目标任务新样本即可推断其是否参与训练
  • 在视觉与语言任务中均成功识别训练包含关系,无需训练数据或影子模型
  • 首次证明仅凭新样本即可完成任务推断,揭示共享表征的隐私漏洞

多任务学习(MTL)通过将多个任务的数据嵌入共同特征空间,实现高效联合训练并减少跨用户数据共享。尽管共享表示被设计为最小必要共享单元,但仍可能无意泄露其训练任务的敏感信息。本文从推理攻击角度研究该问题,提出一种新型黑盒任务推断威胁模型:攻击者仅能查询共享表示对某任务新样本的嵌入向量,目标是判断该任务是否曾用于训练。我们开发了无需影子模型或标注参考数据的纯黑盒攻击方法,利用同一任务内嵌入向量间的依赖关系。在视觉与语言领域多个应用场景中验证,即使仅有来自目标任务分布的新样本,攻击者仍可成功推断任务是否参与训练。理论分析进一步揭示:拥有训练样本的攻击者与仅具新鲜样本的攻击者之间存在严格区分。

原文摘要 · Abstract (English)

Multitask learning (MTL) has emerged as a powerful paradigm that leverages similarities among multiple learning tasks, each with insufficient samples to train a standalone model, to solve them simultaneously while minimizing data sharing across users and organizations. MTL typically accomplishes this goal by learning a shared representation that captures common structure among the tasks by embedding data from all tasks into a common feature space. Despite being designed to be the smallest unit of shared information necessary to effectively learn patterns across multiple tasks, these shared representations can inadvertently leak sensitive information about the particular tasks they were trained on. In this work, we investigate what information is revealed by the shared representations through the lens of inference attacks. Towards this, we propose a novel, black-box task-inference threat model where the adversary, given the embedding vectors produced by querying the shared representation on samples from a particular task, aims to determine whether that task was present when training the shared representation. We develop efficient, purely black-box attacks on machine learning models that exploit the dependencies between embeddings from the same task without requiring shadow models or labeled reference data. We evaluate our attacks across vision and language domains for multiple use cases of MTL and demonstrate that even with access only to fresh task samples rather than training data, a black-box adversary can successfully infer a task's inclusion in training. To complement our experiments, we provide theoretical analysis of a simplified learning setting and show a strict separation between adversaries with training samples and fresh samples from the target task's distribution.

隐私攻击多任务学习黑盒攻击共享表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。