arXiv:2603.04859cs.CRcs.LG2026-03被引 1

用极少污染样本即可劫持模型,揭示合成数据安全风险

Osmosis Distillation: Model Hijacking with the Fewest Samples

  • 仅用少量污染样本生成合成数据,实现模型劫持
  • 攻击成功率高且原任务性能几乎不受影响
  • 适用于多种模型架构,威胁广泛存在

迁移学习利用预训练模型知识,在数据和算力有限的情况下解决新任务。同时,数据蒸馏技术可生成保留原始大数据集关键信息的紧凑数据集。因此,迁移学习与数据蒸馏结合能带来优异性能。然而,使用数据蒸馏生成的合成数据进行迁移学习时,存在显著安全威胁:攻击者仅需在合成数据中注入极少量污染样本,即可实施模型劫持攻击。为揭示此风险,本文提出新型模型劫持策略Osmosis Distillation(OD),仅需最少样本即可完成攻击。在多个数据集上的全面评估表明,该攻击在隐藏任务中成功率高,同时保持原任务的高模型性能。此外,蒸馏得到的渗透数据集可在不同模型架构间实现劫持,使迁移学习中的攻击具有较强性能与实用性。我们强调,使用第三方合成数据进行迁移学习需提高安全意识。

原文摘要 · Abstract (English)

Transfer learning is devised to leverage knowledge from pre-trained models to solve new tasks with limited data and computational resources. Meanwhile, dataset distillation has emerged to synthesize a compact dataset that preserves critical information from the original large dataset. Therefore, a combination of transfer learning and dataset distillation offers promising performance in evaluations. However, a non-negligible security threat remains undiscovered in transfer learning using synthetic datasets generated by dataset distillation methods, where an adversary can perform a model hijacking attack with only a few poisoned samples in the synthetic dataset. To reveal this threat, we propose Osmosis Distillation (OD) attack, a novel model hijacking strategy that targets deep learning models using the fewest samples. Comprehensive evaluations on various datasets demonstrate that the OD attack attains high attack success rates in hidden tasks while preserving high model utility in original tasks. Furthermore, the distilled osmosis set enables model hijacking across diverse model architectures, allowing model hijacking in transfer learning with considerable attack performance and model utility. We argue that awareness of using third-party synthetic datasets in transfer learning must be raised.

模型劫持数据蒸馏安全攻击迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。