提出跨模态对比去学习框架,解决视觉数据删除时的隐私与性能失衡问题。
Preserving Cross-Modal Stability for Visual Unlearning in Multimodal Scenarios
- 用逆对比学习分离视觉表征与原始语义,实现选择性删减
- 保持其他模态判别力,使保留数据类内结构稳定
- 通过双集对比分离扰动,加速去学习且提升准确率7.12%
在自动驾驶等多模态应用中,视觉模态最容易引发隐私泄露;机器去学习旨在从预训练模型中移除特定训练数据以应对隐私风险,但现有方法无法维持跨模态知识和保留数据的类内结构稳定性,导致整体及其它模态性能下降;为此,本文提出跨模态对比去学习(CCU)框架,包含三个核心组件:(a) 选择性视觉去学习:采用逆对比学习将视觉表征与其原始语义解耦;(b) 跨模态知识保留:通过语义一致性保持其他模态的判别能力;(c) 双集对比分离:通过隔离未学习集与保留集之间的结构扰动来维持模型性能;在三个数据集上的大量实验表明,该方法优于现有基线,在仅消耗7%去学习时间下实现7.12%的准确率提升。
原文摘要 · Abstract (English)
Visual modality is the most vulnerable to privacy leakage in real-world multimodal applications like autonomous driving with visual and radar data; Machine unlearning removes specific training data from pre-trained models to address privacy leakage, however, existing methods fail to preserve cross-modal knowledge and maintain intra-class structural stability of retain data, leading to reduced overall and other modalities' performance during visual unlearning; to address these challenges, we propose a Cross-modal Contrastive Unlearning (CCU) framework, which integrates three key components: (a) selective visual unlearning: employing inverse contrastive learning to dissociate visual representations from their original semantics, (b) cross-modal knowledge retention: preserving other modalities' discriminability through semantic consistency, and (c) dual-set contrastive separation: preserving the model performance via isolation of structural perturbations between the unlearn set and retain set; extensive experiments on three datasets demonstrate the superiority of CCU, and our method achieves a 7.12% accuracy improvement with only 7% of the unlearning time compared to the top-accuracy baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。